DNA Codon: Chart, Start and Stop Signals

By Dr. Zubair Khalid, DVM, MS, PhD ·

DNA Codon: Chart, Start and Stop Signals

A codon is a sequence of three consecutive nucleotides in a nucleic acid that specifies one amino acid or a translation stop signal. In DNA, codons are written as triplets of the bases adenine (A), thymine (T), guanine (G) and cytosine (C), and the same information is carried into messenger RNA (mRNA) as triplets of A, uracil (U), G and C after transcription.

Codons matter because they are the physical link between a gene and the protein it encodes. Every protein in a cell, from a bacterial enzyme to a human ion channel, is built by reading codons in order. A single base change inside a codon can swap one amino acid for another, create a premature stop, or leave the protein untouched. That is why the codon table sits at the center of genetics, molecular diagnostics, vaccine design and protein engineering.

What Are Codons in DNA?

The genetic code is a set of rules that maps triplets in messenger RNA to amino acids in proteins [1]. A codon is one entry in that mapping. Because there are four possible bases at each of three positions, the code contains 4 × 4 × 4 = 64 codons.

Three features define how codons behave.

  • Codons are read in non-overlapping triplets. Each base belongs to exactly one codon.
  • Reading starts at a fixed point and continues in one direction, which defines the reading frame.
  • The code is degenerate. Sixty-one codons specify the 20 standard amino acids, and three codons specify stop. Since 61 codons map to 20 amino acids, the mapping cannot be one-to-one, and this redundancy is what biologists call degeneracy [2].

Degeneracy means most amino acids are encoded by more than one codon. Leucine, serine and arginine each have six codons. Methionine and tryptophan have exactly one. Degeneracy is not randomness. It buffers the code against point mutations and shapes how genomes use synonymous codons.

DNA Codons Versus mRNA Codons

A frequent source of confusion is that the codon table is written in RNA letters (U, C, A, G) even though genes are stored in DNA. The table describes the mRNA message, because mRNA is what the ribosome reads. The DNA that produced that mRNA contains the same information in complementary form.

Three sequences are worth keeping separate.

  • The DNA template strand is the strand that is copied during transcription. It is antiparallel to the mRNA and complementary to it, so a template T pairs with mRNA A.
  • The DNA coding strand (also called the nontemplate or sense strand) has the same sequence as the mRNA, except that T replaces U. When people say a gene "contains" a codon, they usually mean the coding strand.
  • The mRNA codon is the triplet actually presented to the ribosome.

For example, if the coding strand reads 5'-ATG-3', the mRNA codon is 5'-AUG-3', and the template strand is 3'-TAC-5'. All three describe the same methionine signal.

The Full 64-Codon Table

The standard genetic code table is conventionally arranged with the first base of the codon down the left side, the second base across the top, and the third base listed within each box. This layout makes the pattern of degeneracy visible: codons that share the first two bases usually encode the same amino acid, and the third base is often interchangeable.

Table 1. The standard genetic code (64 codons). Amino acids are given with their three-letter and one-letter abbreviations. Stop codons are marked with an asterisk.

First baseSecond base USecond base CSecond base ASecond base GThird base
UUUU Phe (F)UCU Ser (S)UAU Tyr (Y)UGU Cys (C)U
UUUC Phe (F)UCC Ser (S)UAC Tyr (Y)UGC Cys (C)C
UUUA Leu (L)UCA Ser (S)UAA Stop *UGA Stop *A
UUUG Leu (L)UCG Ser (S)UAG Stop *UGG Trp (W)G
CCUU Leu (L)CCU Pro (P)CAU His (H)CGU Arg (R)U
CCUC Leu (L)CCC Pro (P)CAC His (H)CGC Arg (R)C
CCUA Leu (L)CCA Pro (P)CAA Gln (Q)CGA Arg (R)A
CCUG Leu (L)CCG Pro (P)CAG Gln (Q)CGG Arg (R)G
AAUU Ile (I)ACU Thr (T)AAU Asn (N)AGU Ser (S)U
AAUC Ile (I)ACC Thr (T)AAC Asn (N)AGC Ser (S)C
AAUA Ile (I)ACA Thr (T)AAA Lys (K)AGA Arg (R)A
AAUG Met (M)ACG Thr (T)AAG Lys (K)AGG Arg (R)G
GGUU Val (V)GCU Ala (A)GAU Asp (D)GGU Gly (G)U
GGUC Val (V)GCC Ala (A)GAC Asp (D)GGC Gly (G)C
GGUA Val (V)GCA Ala (A)GAA Glu (E)GGA Gly (G)A
GGUG Val (V)GCG Ala (A)GAG Glu (E)GGG Gly (G)G

How to Read the Table

To look up a codon, read the three bases in order and find the row matching the first base, the column matching the second base, and the third-base line inside that box. For example, the codon 5'-AUG-3' sits in the A row, the U column, on the G line, and it reads Met (M). The codon 5'-UGG-3' sits in the U row, the G column, on the G line, and it reads Trp (W).

The table also reveals a structural pattern that has attracted decades of theoretical work. The arrangement is not arbitrary. Symmetry analyses of the code show a purine-pyrimidine symmetry net that is shared across nuclear and mitochondrial codes, and this net is not visible in the alphabetical U-C-A-G layout of the standard table [3][4][5]. These symmetry studies treat the third base as functionally dominant in some respects, which contrasts with the traditional view that the third position is mostly silent [4].

Degeneracy and the Wobble Position

The third base of a codon is called the wobble position. Francis Crick proposed in 1966 that base pairing between the third codon base and the first anticodon base is less stringent than pairing at the other two positions, which allows one transfer RNA (tRNA) to read more than one codon. This is the wobble hypothesis, and it explains why codons ending in different bases can specify the same amino acid.

Wobble is not a free-for-all. Computational studies of RNA base pair geometry show that hydrogen-bonding configurations and base pair width screen which pairs are allowed at the wobble position, with guanine, hypoxanthine and queuine behaving as predictable anticodon wobble bases [6][7]. The exclusion of adenine from the anticodon wobble position cannot be explained by pairing geometry alone, which suggests that other factors, including base-sugar interactions, contribute [6].

Post-transcriptional modifications at tRNA position 34, the wobble nucleotide, and position 37, adjacent to the anticodon, expand the reading capacity of the code. In Escherichia coli, the full set of tRNA modifications and the genes that install them has been mapped, and integrating that modification pattern into the codon table produces a practical reference for predicting how each codon is read [8]. Structural work on the 70S ribosome showed that the hypermodified base mnm⁵s²U in tRNA^Lys reads both AAA and AAG and discriminates against near-cognate codons such as the stop codon UAA, using shape complementarity rather than simple hydrogen bonding [9].

Wobble also has consequences for codon usage. Analysis of MS2 phage RNA found a bias favoring cytidine over uracil among codons that require a third-position pyrimidine, which is consistent with selection against wobble pairing in the tRNA-mRNA interaction [10]. Among fourfold degenerate codons, pyrimidines are favored over purines [10].

Start and Stop Signals

Translation needs clear boundaries. The start codon marks where protein synthesis begins, and stop codons mark where it ends.

The Start Codon

AUG is the start codon in the standard genetic code. It encodes methionine, and in most mRNAs it is the first codon translated. In bacteria the initiator is a modified formylmethionine, but the codon is still AUG. Rare alternative start codons such as GUG and UUG can be used in bacteria, but AUG is the canonical signal.

Start codon selection is regulated, not automatic. Upstream open reading frames (uORFs) in the 5' untranslated region can capture the ribosome before it reaches the main coding sequence. In the KCNQ2 gene, which encodes a neuronal voltage-gated potassium channel, a single uORF is strongly repressive of translation, and mutations that disable the uORF start codon enhance synthesis of the encoded channel [11]. Adenine base editing of that uORF start codon weakens ribosome engagement at the uORF and increases translation of the downstream protein in a neuron-like cell line [11]. This is a direct demonstration that the start codon is a control point, not just a punctuation mark.

The Stop Codons

Three codons do not encode amino acids. UAA, UAG and UGA are stop codons, also called nonsense codons or termination codons. When the ribosome reaches one of them, no matching tRNA delivers an amino acid, and the completed polypeptide is released.

The three stop codons are sometimes called amber (UAG), ochre (UAA) and opal (UGA). Their sequences are fixed in the standard code, but their meaning is not universal across all genetic systems, which is discussed below.

Reading Frames: Why the Starting Point Matters

A reading frame is the grouping of nucleotides into triplets that the ribosome uses. Because codons are three bases long, any stretch of RNA can be read in three different frames depending on where you start.

Consider the sequence 5'-AUGGCUUAA-3'. Read from the first base, the frame is AUG (Met), GCU (Ala), UAA (Stop). Shift the start by one base and the frame becomes UGG (Trp), CUU (Leu), and a trailing A that cannot form a complete codon. Shift by two and you get GGC (Gly), UUA (Leu).

Only one of these frames produces the intended protein. The others produce unrelated amino acid sequences, and they usually encounter a stop codon quickly. This is why insertions or deletions that are not multiples of three are so damaging. They shift the frame for the entire downstream message, a class of mutation called a frameshift.

Frameshifting is also used deliberately. Picornaviruses, which have single-stranded positive-sense RNA genomes, were long thought to contain a single open reading frame preceded by a long 5' untranslated region with an internal ribosomal entry site. Recent work shows that many picornavirus genera carry an additional open reading frame that overlaps the 5' end of the main one, or lies entirely within it in a different reading frame, and that access to these alternate frames depends on mechanisms including leaky scanning, termination-reinitiation and ribosomal frameshifting [12]. The same triplet sequence can therefore encode two different proteins depending on the frame the ribosome selects.

Worked Example: Translating a Short mRNA

This example walks a full message through translation, one codon at a time.

mRNA: 5'-AUG GCC UAC GGG UUU UAA-3'

Step 1. Identify the reading frame. The first AUG from the 5' end sets the frame. Everything is read in triplets from that point.

Step 2. Read each codon against the table.

PositionCodonAmino acidNotes
1AUGMet (M)Start codon
2GCCAla (A)
3UACTyr (Y)
4GGGGly (G)
5UUUPhe (F)
6UAAStopTermination

Step 3. Assemble the peptide. The growing chain is Met-Ala-Tyr-Gly-Phe. Each amino acid is added to the carboxyl end of the previous one, so the sequence is written from the amino terminus to the carboxyl terminus.

Step 4. Terminate. When the ribosome reaches UAA, no tRNA with a matching anticodon enters the A site. Release factors recognize the stop codon, the polypeptide is released, and the ribosome dissociates.

Step 5. Trace back to DNA. If the coding strand of the gene reads 5'-ATG GCC TAC GGG TTT TAA-3', the template strand is 3'-TAC CGG ATG CCC AAA ATT-5'. Transcription of the template produces the mRNA above, with U replacing T.

This example shows the full logic of the code in one pass. The start codon fixes the frame, the table assigns amino acids, and the stop codon ends the message.

How the Code Is Observed in Practice

Codon-level information is read directly from sequence data in routine laboratory work.

  • Sanger sequencing of a cloned gene returns the DNA sequence, and the coding strand is translated in silico using the standard table to predict the protein.
  • Next-generation sequencing of an mRNA population (RNA-seq) identifies which codons are present and how often each is used.
  • Variant annotation tools classify a single-base change as synonymous, missense or nonsense based on the codon it alters. A nonsense variant creates a premature stop codon and usually truncates the protein.
  • Base editing and prime editing are designed at the codon level. The KCNQ2 uORF work used adenine base editing to alter a start codon and change translation output, which is a direct example of codon-level therapeutic design [11].

For any of these workflows, the reference point is a translation table. The National Center for Biotechnology Information maintains a set of genetic code tables used for sequence annotation, and the standard code is table 1.

Comparative and Clinical Relevance

The standard genetic code is nearly universal. The same 64 codons specify the same amino acids across bacteria, archaea, plants and animals, which is why a human gene can often be expressed in E. coli with only codon-usage adjustments. Symmetry analyses of the code support this deep conservation. The purine-pyrimidine symmetry net that underlies the code is described as common to more than 30 known genetic codes, including nuclear and mitochondrial variants, and it remains unchanged across evolution [3][4][5][13].

Universality is not absolute. Mitochondria use variant codes. In vertebrate mitochondrial code, UGA encodes tryptophan rather than stop, and AUA encodes methionine rather than isoleucine. Other variant codes exist in ciliates, in some fungi and in certain archaea. The NCBI translation table collection lists these variants separately, and any annotation pipeline must select the correct table for the organism.

Codon-level changes drive human disease. A missense variant that swaps a polar residue for a hydrophobic one can misfold a protein. A nonsense variant that creates a premature stop triggers nonsense-mediated decay or produces a truncated protein. A frameshift from a single-base insertion changes every downstream codon. In each case, the clinical interpretation depends on reading the codon table correctly.

Codon degeneracy also has measurable biological consequences beyond protein sequence. Mathematical modeling of degeneracy using nucleotide base classifications and Hamming distances has been used to characterize differences between gram-positive and gram-negative bacterial genes, suggesting that codon bias carries information about organism-level behavior [2]. Codon reassignment events across the 27 NCBI translation tables tend to preserve codon-family connectivity, which suggests that natural code variants are constrained by the same error-minimizing logic that shapes the standard code [14].

RNA editing adds another layer. Adenosine-to-inosine editing can recode a codon at the mRNA level without changing the DNA. In the well-studied mammalian Gln-to-Arg recoding site, the ancestral state in vertebrate genomes is the pre-editing Gln codon, and all 470 available mammalian genomes avoid the other three equivalent ways to achieve Arg in protein, which points to strong selection on the editing motif and structure rather than on the codon alone [15].

Common Mistakes and Limitations

Reading the table in DNA letters. The standard table uses U, not T. If you are translating a DNA coding strand, substitute T with U before looking up codons, or use a DNA-equivalent table.

Confusing the template and coding strands. The template strand is complementary to the mRNA. The coding strand matches the mRNA except for T and U. Mixing them up reverses every codon.

Assuming the first AUG is always the start. In eukaryotes, the ribosome usually scans from the 5' cap and initiates at the first AUG in a favorable context, but uORFs and context effects mean the first AUG is not always used. In the KCNQ2 example, an upstream AUG represses translation of the main protein [11].

Treating the third base as meaningless. Wobble makes the third base flexible, but it is not free. tRNA modifications, codon usage bias and pairing geometry all constrain which synonymous codons are used efficiently [6][7][8][9][10].

Forgetting that the code varies. Mitochondrial and several nuclear variant codes reassign codons. Using the standard table for a mitochondrial gene will produce the wrong protein.

Ignoring the reading frame. A sequence can be translated in three frames. Only one is correct. Frameshift mutations and overlapping open reading frames both exploit this [12].

Assuming degeneracy is pure redundancy. Degeneracy affects translation speed, mRNA stability and folding. Synonymous codons are not always interchangeable in practice.

Individual patient or animal cases require professional evaluation. The codon table tells you what a sequence encodes, not what it means clinically in a specific organism.

Quick Review

  • A codon is three consecutive nucleotides that specify one amino acid or a stop signal.
  • The code has 64 codons: 61 sense codons for 20 amino acids and 3 stop codons.
  • AUG is the start codon and encodes methionine.
  • UAA, UAG and UGA are the three stop codons.
  • Codons are read in non-overlapping triplets from a fixed reading frame.
  • Degeneracy means most amino acids have more than one codon, and the third (wobble) position is the main source of flexibility.
  • The code is nearly universal, but mitochondrial and other variant codes reassign specific codons.

Frequently Asked Questions

What are codons in DNA?

Codons in DNA are triplets of the bases A, T, G and C that correspond to the codons in mRNA after transcription. The DNA coding strand carries the same information as the mRNA, with T in place of U.

How many codons are there?

There are 64 codons. Sixty-one specify amino acids and three (UAA, UAG, UGA) specify stop.

Which codon starts translation?

AUG is the start codon. It encodes methionine and sets the reading frame for the rest of the message.

What are the three stop codons?

UAA, UAG and UGA. They are also called ochre, amber and opal, and they terminate translation by recruiting release factors instead of tRNA.

Why is the genetic code called degenerate?

Because 61 codons encode only 20 amino acids, most amino acids are specified by more than one codon. This redundancy is called degeneracy [2].

Is the genetic code the same in all organisms?

Nearly, but not exactly. Mitochondria and some nuclear lineages use variant codes. For example, vertebrate mitochondria read UGA as tryptophan and AUA as methionine.

Related Articles

Sources

  1. Relational model of the standard genetic code.
  2. Analysing the genetic code degeneracy: a consequence towards bacterial staining.
  3. The Supersymmetry Genetic Code Table and Quadruplet Symmetries of DNA Molecules Are Unchangeable and Synchronized with Codon-Free Energy Mapping during Evolution.
  4. Standard Genetic Code vs. Supersymmetry Genetic Code - Alphabetical table vs. physicochemical table.
  5. Maximal Genetic Code Symmetry Is a Physicochemical Purine-Pyrimidine Symmetry Language for Transcription and Translation in the Flow of Genetic Information from DNA to Proteins.
  6. Role of wobble base pair geometry for codon degeneracy: purine-type bases at the anticodon wobble position.
  7. Configuration of wobble base pairs having pyrimidines as anticodon wobble bases: significance for codon degeneracy.
  8. A tRNA modification pattern that facilitates interpretation of the genetic code.
  9. Novel base-pairing interactions at the tRNA wobble position crucial for accurate reading of the genetic code.
  10. Is there selection against wobble in codon-anticodon pairing?
  11. An upstream open reading frame represses translation of the neuronal potassium channel KCNQ2.
  12. The Dicistronic Nature of Picornavirus Genomes? Diverse Translation Mechanisms Enable the Expression of Accessory Proteins from Alternate Open Reading Frames.
  13. The Evolution of Life Is a Road Paved with the DNA Quadruplet Symmetry and the Supersymmetry Genetic Code.
  14. Robust error-minimization in the genetic code across physicochemical metrics and variant codes: A graph-theoretic analysis in GF(2)(6).
  15. Learning from the Codon Table: Convergent Recoding Provides Novel Understanding on the Evolution of A-to-I RNA Editing.