Codon Table: How to Read the Genetic Code for Translation

By Dr. Zubair Khalid, DVM, MS, PhD ·

Codon Table: How to Read the Genetic Code for Translation

Introduction to the Codon Table

The codon table is the reference chart that maps each three-nucleotide sequence of messenger RNA (mRNA) to its corresponding amino acid or a translational signal. This mapping constitutes the genetic code, the universal language by which nucleic acid sequence information is translated into protein sequence. Without the codon table, the central dogma of molecular biology—DNA to RNA to protein—would be unreadable.

The table is not merely a pedagogical convenience; it is the product of decades of experimental work that determined which of the 64 possible triplet combinations correspond to each of the 20 standard amino acids, the start signal, and the three stop signals. For any student of molecular biology, mastering the codon table is non-negotiable. It is the tool you will use to translate mRNA sequences into protein sequences, to predict the effects of mutations, and to design experiments involving gene expression. This article provides a complete, mechanistic guide to the codon table: what it is, how to read it, and why it behaves the way it does.

The Genetic Code: Codons and Amino Acids

Triplet Code

The genetic code is read in units of three nucleotides called codons. Each codon specifies a single amino acid. The triplet nature of the code was established theoretically by George Gamow in 1954 and confirmed experimentally by Francis Crick and Sydney Brenner in 1961 using proflavin-induced frameshift mutations in the rII locus of bacteriophage T4. A single nucleotide insertion or deletion shifts the reading frame, producing a completely different protein sequence downstream of the mutation. Two insertions or two deletions also disrupt the frame, but three insertions or three deletions restore the correct reading frame, leaving only a small insertion or deletion of amino acids. This genetic proof demonstrated that the code is read in non-overlapping triplets.

The reading frame is established at the start of translation and proceeds in the 5′ to 3′ direction along the mRNA. There is no punctuation between codons; the ribosome simply moves three nucleotides at a time. This means the same mRNA sequence can theoretically be read in three different frames, but only one frame produces a functional protein in vivo, typically the one beginning at the Start Codon.

Start and Stop Codons

Of the 64 possible codons, 61 are sense codons that specify amino acids. The remaining three—UAA, UAG, and UGA—are nonsense codons or Stop Codons. They do not code for any amino acid; instead, they signal the ribosome to terminate translation and release the completed polypeptide.

One codon, AUG, serves a dual function. It codes for the amino acid methionine, and it also serves as the primary Starting Codon that initiates translation. In bacteria, the initiating AUG is recognized by a specialized initiator tRNA carrying formyl-methionine (fMet), while in eukaryotes, the initiating AUG is recognized by the initiator tRNA carrying unmodified methionine. In both domains, internal AUG codons code for standard methionine.

The 20 standard amino acids are encoded by the 61 sense codons, but the distribution is not uniform. Methionine and tryptophan are each encoded by a single codon (AUG and UGG, respectively). All other amino acids are encoded by two, three, four, or six codons. Leucine, arginine, and serine are each encoded by six codons, the maximum degeneracy in the standard code.

How to Read a Codon Table

Reading the Rows and Columns

A standard codon table is organized as a grid with three axes: the first nucleotide (5′ base), the second nucleotide (middle base), and the third nucleotide (3′ base). The table is typically arranged with the first base along the left margin, the second base along the top, and the third base along the right margin.

To read the table, follow these steps:

  1. Identify the first nucleotide of the codon (the 5′ base). Locate it in the leftmost column of the table.
  2. Identify the second nucleotide (the middle base). Locate it in the top row of the table.
  3. Find the cell at the intersection of the first-base row and the second-base column. This cell contains a list of amino acids, one for each possible third base.
  4. Identify the third nucleotide (the 3′ base). Locate it in the rightmost column of the table, within the row of the cell you found in step 3.
  5. Read the amino acid at the intersection of the third base and the cell.

For example, consider the codon GCA. The first base is G, the second base is C, and the third base is A. Locate the G row and the C column. The cell at this intersection contains four amino acids: alanine (for GCU, GCC, GCA, and GCG). Reading across the third base column, the A entry gives alanine. Therefore, GCA codes for alanine.

Example: Decoding AUG

Let us decode AUG step by step:

  1. First base: A (adenine). Locate the A row.
  2. Second base: U (uracil). Locate the U column.
  3. The intersection of the A row and U column contains four entries: AUU (isoleucine), AUC (isoleucine), AUA (isoleucine), and AUG (methionine).
  4. Third base: G (guanine). The G entry in this cell is methionine.
  5. Therefore, AUG codes for methionine and serves as the Start Codon.

This same procedure applies to all 64 codons. With practice, you will memorize the most common codons, but the table is always available as a reference.

The Standard Codon Table Diagram

RNA vs DNA Tables

Codon tables are almost always presented using RNA bases (A, U, G, C) because translation occurs on mRNA. However, some textbooks and problem sets present DNA-based tables using T instead of U. When using a DNA table, you are reading the coding strand of DNA, which has the same sequence as the mRNA except that thymine (T) replaces uracil (U). To convert between the two, simply substitute T for U or U for T.

It is critical to know which type of table you are using. If you are given a DNA template strand sequence and asked to predict the protein, you must first transcribe it to mRNA (replacing T with U and using the complementary sequence) before consulting an RNA codon table. If you are given a coding strand DNA sequence, you can use a DNA codon table directly, but you must remember that the protein sequence will be identical to what an RNA table would give after substituting U for T.

Color-Coded Amino Acid Groups

Most codon table diagrams color-code amino acids by their chemical properties. This is not arbitrary decoration; it highlights patterns in the genetic code that reflect the physicochemical relationships between codons and amino acids.

  • Nonpolar (hydrophobic) amino acids (glycine, alanine, valine, leucine, isoleucine, methionine, phenylalanine, tryptophan, proline) are often shown in one color.
  • Polar uncharged amino acids (serine, threonine, cysteine, tyrosine, asparagine, glutamine) are shown in another.
  • Positively charged amino acids (lysine, arginine, histidine) are shown in a third.
  • Negatively charged amino acids (aspartate, glutamate) are shown in a fourth.

These groupings reveal a key feature of the code: codons sharing the same first two bases often encode amino acids with similar chemical properties. For example, all codons beginning with GU (GUU, GUC, GUA, GUG) encode valine, a hydrophobic amino acid. All codons beginning with GA (GAU, GAC, GAA, GAG) encode aspartate or glutamate, both negatively charged. This pattern minimizes the deleterious effects of point mutations, particularly at the third base position, a property known as the "error minimization" of the genetic code.

Degeneracy and Wobble in the Genetic Code

Synonymous Codons

The genetic code is degenerate: most amino acids are encoded by more than one codon. Codons that specify the same amino acid are called synonymous codons. For example, leucine is encoded by UUA, UUG, CUU, CUC, CUA, and CUG. Synonymous codons are not used with equal frequency; organisms exhibit codon usage bias, favoring certain synonymous codons over others based on tRNA abundance and translational efficiency. Highly expressed genes in Escherichia coli, for instance, preferentially use codons that match the most abundant tRNAs.

Degeneracy provides a buffer against mutation. A single nucleotide substitution at the third position of a codon often results in a silent mutation—a change in the DNA sequence that does not alter the amino acid sequence. This is because many synonymous codons differ only at the third base. For example, GGU, GGC, GGA, and GGG all code for glycine; a change at the third position from U to C, A, or G has no effect on the protein product.

Third Base Wobble

The degeneracy of the genetic code is made possible by wobble, a phenomenon first described by Francis Crick in 1966. The interaction between a codon on mRNA and the complementary anticodon on tRNA follows standard Watson-Crick base pairing (A-U, G-C) at the first two positions of the codon. However, the third position of the codon (the 3′ base) can form non-standard base pairs with the first position of the anticodon (the 5′ base).

The wobble rules are:

  • The anticodon base inosine (I), which is a modified base found in tRNA, can pair with U, C, or A in the third codon position.
  • The anticodon base U can pair with A or G in the third codon position.
  • The anticodon base G can pair with U or C in the third codon position.

This flexibility means that a single tRNA can recognize multiple codons. For example, a tRNA with the anticodon 3′-UAI-5′ (where I is inosine) can recognize the codons 5′-AUU-3′, 5′-AUC-3′, and 5′-AUA-3′, all of which code for isoleucine. Wobble reduces the number of tRNAs required: most organisms have roughly 30–40 tRNA species, far fewer than the 61 sense codons would require if each codon needed a dedicated tRNA. For a detailed discussion of the codon–anticodon interaction, see Codon Anticodon.

Start and Stop Codons: Initiating and Terminating Translation

Alternative Start Codons

While AUG is the canonical start codon, it is not the only one. In bacteria, GUG and UUG can also serve as start codons, though they are less efficient. When GUG or UUG is used as a start codon, the initiating tRNA still carries methionine (or formyl-methionine in bacteria), not valine or leucine. The initiator tRNA recognizes the start codon through its anticodon (3′-UAC-5′), which pairs with AUG. For GUG and UUG, the pairing is less precise, relying on wobble at the first position of the codon, which is why these alternative start codons are used less frequently.

In eukaryotes, AUG is overwhelmingly the dominant start codon, and the context surrounding it matters. The Kozak consensus sequence (gccRccAUGG, where R is a purine) in vertebrates enhances initiation efficiency. When the ribosome scans the mRNA from the 5′ cap, it typically initiates at the first AUG that lies in a favorable context.

The Start Codon AUG is also the codon for methionine, meaning that all newly synthesized proteins in eukaryotes begin with methionine. In bacteria, the N-terminal amino acid is formyl-methionine, which is often removed post-translationally by deformylase and methionine aminopeptidase.

Nonsense Mutations

A nonsense mutation is a point mutation that changes a sense codon into a Stop Codon (UAA, UAG, or UGA). This causes premature termination of translation, producing a truncated protein that is usually nonfunctional. Nonsense mutations are often highly deleterious. For example, in the CFTR gene, the nonsense mutation G542X (a G-to-T change at codon 542, converting GGA to TGA in the DNA, which is UGA in mRNA) causes cystic fibrosis by producing a truncated CFTR protein.

The three stop codons are also called termination codons or Termination Codon. They are recognized not by tRNAs but by release factors: RF1 and RF2 in bacteria (RF1 recognizes UAA and UAG; RF2 recognizes UAA and UGA) and eRF1 in eukaryotes (recognizes all three stop codons). The release factors trigger hydrolysis of the peptidyl-tRNA bond, releasing the completed polypeptide from the ribosome.

Experimental Methods Used to Decipher the Codon Table

Poly-U Experiment

The first codon to be deciphered was UUU, coding for phenylalanine. This was achieved in 1961 by Marshall Nirenberg and Heinrich Matthaei. They synthesized a simple mRNA consisting solely of uracil residues (poly-U) and added it to a cell-free translation system derived from E. coli that contained ribosomes, tRNAs, amino acids, and the necessary translation factors. They included one radioactively labeled amino acid at a time in separate reactions. Only when phenylalanine was the labeled amino acid did they observe the synthesis of a polypeptide (polyphenylalanine). This demonstrated that UUU codes for phenylalanine.

The same approach was extended to poly-A (coding for lysine) and poly-C (coding for proline). Poly-G could not be tested easily because poly-G RNA forms complex secondary structures that do not translate efficiently.

Ribosome Binding Assay

The poly-U experiment could only reveal the codons for homopolymers. To decode mixed codons, Nirenberg and Philip Leder developed the triplet binding assay in 1964. They found that a trinucleotide (a codon of exactly three nucleotides) could promote the binding of a specific aminoacyl-tRNA to a ribosome, even in the absence of translation. The ribosome–tRNA–trinucleotide complex could be trapped on a nitrocellulose filter, allowing the researchers to determine which aminoacyl-tRNA bound to which trinucleotide.

For example, when the trinucleotide UUU was incubated with ribosomes and a mixture of aminoacyl-tRNAs, only phenylalanyl-tRNA bound. By systematically testing all 64 trinucleotides, Nirenberg and Leder assigned amino acids to most codons. The remaining assignments were filled in using repeating copolymers (e.g., poly-UG with a repeating UGUGUG sequence) and other techniques. By 1966, the complete genetic code was known.

Common Pitfalls and Misconceptions

Directionality (5' to 3')

The most common error students make when reading a codon table is ignoring directionality. Codons are always written and read in the 5′ to 3′ direction. The first base of the codon is the 5′ base, and the third base is the 3′ base. If you write a codon backwards (e.g., reading ACG as GCA), you will get the wrong amino acid. For example, ACG codes for threonine, while GCA codes for alanine.

When you are given a DNA template strand, you must remember that the mRNA is synthesized antiparallel to the template. If the template strand is 3′-TAC-5′, the mRNA will be 5′-AUG-3′, which codes for methionine. A common mistake is to read the template strand directly as if it were the coding strand.

Distinguishing T and U

Another frequent error is confusing thymine (T) and uracil (U). The codon table is RNA-based, so it uses U. If you are working with DNA sequences, you must convert T to U before using the table. Conversely, if you are asked to write a DNA sequence corresponding to a given protein, you must convert U to T.

A related issue arises with DNA codon tables. Some tables present the coding strand of DNA (with T instead of U). If you accidentally use an RNA table with a DNA sequence, you will misread every codon containing T. Always check which nucleic acid the table is based on before you begin.

Misreading the Third Base

Because of wobble, the third base of a codon is often the least important for determining the amino acid. However, this does not mean you can ignore it. The third base distinguishes between codons that code for different amino acids in some cases. For example, CAU and CAC both code for histidine, but CAA and CAG code for glutamine. If you ignore the third base, you will confuse histidine and glutamine.

Assuming the Table Is Universal Without Exception

The standard genetic code is nearly universal, but it is not absolutely universal. Mitochondrial genomes and some ciliates use variant codes. For example, in human mitochondria, UGA codes for tryptophan instead of being a stop codon, and AUA codes for methionine instead of isoleucine. For standard undergraduate problems, the universal code applies, but you should be aware that exceptions exist.

Practical Summary: Using the Codon Table in Problem Solving

Step-by-Step Translation

To translate an mRNA sequence into a protein sequence, follow these steps:

  1. Write the mRNA sequence in the 5′ to 3′ direction. If you are given a DNA template strand, first synthesize the complementary mRNA (A pairs with U, T pairs with A, C pairs with G, G pairs with C).
  2. Locate the start codon. Find the first AUG in the sequence. Translation begins at this codon.
  3. Group the sequence into codons of three nucleotides, starting at the AUG. Do not skip any nucleotides.
  4. Read each codon using the codon table, writing down the corresponding amino acid.
  5. Stop at the first stop codon (UAA, UAG, or UGA). Do not include the stop codon in the protein sequence.
  6. Write the protein sequence using the one-letter or three-letter amino acid abbreviations.

Practice Example

Translate the following mRNA sequence:

5′-AUGGCUAAAUGCUGA-3′

  1. Start codon: AUG at position 1.
  2. Group into codons: AUG | GCU | AAA | UGC | UGA
  3. Read each codon:
  4. AUG = Methionine (Met, M)
  5. GCU = Alanine (Ala, A)
  6. AAA = Lysine (Lys, K)
  7. UGC = Cysteine (Cys, C)
  8. UGA = Stop
  9. Protein sequence: Met-Ala-Lys-Cys (MAKC)

The stop codon UGA is not translated into an amino acid; it terminates translation.

For a visual alternative to the table, you may find the Codon Wheel useful. The wheel presents the same information in a circular format, which some students find faster to use once they are familiar with it.

Frequently Asked Questions

How do you read a codon table?

To read a codon table, identify the first nucleotide of the codon in the leftmost column, the second nucleotide in the top row, and the third nucleotide in the rightmost column. The amino acid is found at the intersection of these three positions. Always read codons in the 5′ to 3′ direction.

What is the difference between a codon table and an amino acid chart?

A codon table maps mRNA codons (three-nucleotide sequences) to amino acids. An amino acid chart typically lists the 20 amino acids with their structures, properties, and abbreviations. Some amino acid charts include codon information, but their primary purpose is to summarize amino acid chemistry, not to decode mRNA.

Why are there 64 codons but only 20 amino acids?

There are 64 possible triplet combinations of four nucleotides (4³ = 64). Only 20 standard amino acids exist, so the code is degenerate: most amino acids are encoded by more than one codon. Three codons are stop signals, and one codon (AUG) serves as both the start signal and the codon for methionine.

What does the start codon AUG code for?

AUG codes for methionine. In bacteria, the initiating AUG carries formyl-methionine (fMet), a modified form of methionine. In eukaryotes, the initiating AUG carries standard methionine. Internal AUG codons also code for methionine.

What are stop codons and what do they do?

Stop codons are UAA, UAG, and UGA. They do not code for any amino acid. Instead, they are recognized by release factors that trigger the hydrolysis of the peptidyl-tRNA bond, releasing the completed polypeptide from the ribosome and terminating translation.

Is the codon table the same for all organisms?

The standard genetic code is nearly universal, but exceptions exist. Mitochondrial genomes and some ciliates use variant codes. For example, in vertebrate mitochondria, UGA codes for tryptophan and AUA codes for methionine. For standard coursework, the universal code is used unless stated otherwise.

How do you remember the codon table?

Memorization strategies vary, but a few patterns help. The codons for a given amino acid often share their first two bases. For example, all CCN codons (CCU, CCC, CCA, CCG) code for proline. The third base is often the least important. You can also memorize the single-codon amino acids (methionine = AUG, tryptophan = UGG) and the stop codons (UAA, UAG, UGA) first, then learn the rest by practice. Repeated translation problems are the most effective way to internalize the table.

Key Takeaways

  • The codon table maps 64 mRNA triplets to 20 amino acids, one start signal, and three stop signals.
  • Codons are read in the 5′ to 3′ direction, and the table is RNA-based (using U, not T).
  • AUG is the primary start codon and codes for methionine; UAA, UAG, and UGA are stop codons that terminate translation.
  • The genetic code is degenerate: most amino acids are encoded by multiple codons, which provides robustness against mutations.
  • Wobble at the third codon position allows a single tRNA to recognize multiple codons, reducing the number of tRNAs required.
  • The code was deciphered through the poly-U experiment and the triplet binding assay by Nirenberg, Matthaei, and Leder.
  • When translating mRNA, locate the start codon, group nucleotides into triplets, read each codon, and stop at the first stop codon.

Further Reading

  • Mohanta TK et al. Construction of anti-codon table of the plant kingdom and evolution of tRNA selenocysteine (tRNA(Sec)). BMC genomics. 2020. PubMed 33213362
  • Ma L et al. Learning from the Codon Table: Convergent Recoding Provides Novel Understanding on the Evolution of A-to-I RNA Editing. Journal of molecular evolution. 2024. PubMed 39012510
  • Serra N, Di Carlo P. Codon and Reverse Codon: A Theoretical Approach to Reinterpret the Genetic Code Table. Cureus. 2023. PubMed 38084181
  • Cherry JM. Codon usage table for Xenopus laevis. Methods in cell biology. 1991. PubMed 181115960304-0)
  • Rosandić M, Paar V. The Supersymmetry Genetic Code Table and Quadruplet Symmetries of DNA Molecules Are Unchangeable and Synchronized with Codon-Free Energy Mapping during Evolution. Genes. 2023. PubMed 38137022

Related Clinical & Scientific Guides