Starting Codon: Initiation of Protein Synthesis Explained

By Dr. Zubair Khalid, DVM, MS, PhD ·

Starting Codon: Initiation of Protein Synthesis Explained

The starting codon is the first codon of a messenger RNA (mRNA) transcript that is translated into protein. In the vast majority of organisms, this codon is AUG, which specifies the amino acid methionine. The starting codon defines the reading frame for the entire downstream coding sequence, establishing which of the three possible triplet groupings of nucleotides will be read as codons. A shift of even one nucleotide in this frame would produce a completely different, and almost certainly non-functional, protein. Because translation proceeds in the 5′ to 3′ direction along the mRNA, the starting codon is always the AUG triplet closest to the 5′ end of the coding region, though as we will see, context matters considerably.

Definition and Basic Concept

A codon is a sequence of three nucleotides that corresponds to a specific amino acid or a translational signal. The starting codon, also called the initiation codon, is the specific codon at which the ribosome assembles and begins protein synthesis. In the standard genetic code, AUG is the sole start codon, and it simultaneously codes for methionine. This dual function—serving both as a "start here" signal and as an amino acid specification—is a fundamental feature of translation.

The relationship between codons and amino acids is described by the Codon Definition, and the complete set of 64 possible triplet combinations is organized in the Codon Table. Of these 64 codons, 61 code for amino acids and 3 are Stop Codon signals (UAA, UAG, UGA) that terminate translation. AUG is unique in that it is the only codon that serves as both an amino acid codon (methionine) and a universal start signal in the nuclear genomes of eukaryotes and prokaryotes.

The starting codon sets the reading frame. Consider an mRNA sequence: 5′-AUGGCUAAACGA-3′. The ribosome begins at AUG and reads successive triplets: AUG (Met), GCU (Ala), AAA (Lys), CGA (Arg). If translation were to begin one nucleotide downstream at UGG, the reading frame would shift entirely, producing a different amino acid sequence. This frame-setting function is arguably more important than the methionine specification itself, as it ensures the correct linear order of amino acids in the nascent polypeptide.

The Genetic Code and Start Codons

The genetic code is the set of rules by which nucleotide triplets in mRNA specify amino acids in proteins. It is nearly universal across all life forms, from bacteria to humans, which strongly suggests that the code was fixed early in evolutionary history. The code is degenerate, meaning that most amino acids are specified by more than one codon. Methionine and tryptophan are the only amino acids with a single codon each (AUG and UGG, respectively).

Universality and Exceptions

The use of AUG as the start codon is nearly universal, but there are notable exceptions. Mitochondria, which possess their own small genomes and translational machinery, show the most variation. In human mitochondria, the codons AUA and AUU can also serve as start codons, and mitochondrial initiation does not always require methionine. In some mitochondrial genomes, the start codon can be GUG or UUG, which normally code for valine and leucine, respectively, in the standard code. These variations are possible because mitochondrial translation systems have simplified initiation requirements and a reduced set of transfer RNAs (tRNAs).

Some prokaryotes and bacteriophages also use alternative start codons. For example, Escherichia coli uses GUG and UUG as start codons in roughly 8% and 1% of its genes, respectively. When GUG or UUG serves as a start codon, the initiating amino acid is still methionine, not valine or leucine, because the initiator tRNA charged with methionine can base-pair with these codons in the ribosomal P site with reduced fidelity. This is a critical point: the start codon identity and the amino acid identity are decoupled at initiation.

Alternative Start Codons

Alternative start codons are codons that can initiate translation but are not AUG. In addition to GUG and UUG, CUG is used in some yeast species and in certain mammalian mRNAs. In the yeast Saccharomyces cerevisiae, CUG is a rare start codon, and in some mammalian oncogenes, CUG-initiated translation produces N-terminally extended protein isoforms with distinct functions. The efficiency of initiation at non-AUG codons is generally lower than at AUG, typically 1–10% of the AUG efficiency, which provides a mechanism for producing low levels of alternative protein isoforms from the same mRNA.

The Codon Wheel is a useful visual tool for understanding how codons are organized and how single-nucleotide changes can alter codon identity. It is particularly helpful when considering how mutations might create or destroy start codons.

Role in Translation Initiation

Translation initiation is the most complex and highly regulated step of protein synthesis. It requires the coordinated action of the ribosome, mRNA, initiator tRNA, and multiple initiation factors. The starting codon is the focal point of this process, as it is the site where the small ribosomal subunit first establishes a stable interaction with the mRNA.

Initiation Complex Formation

In prokaryotes, initiation proceeds as follows:

  1. The small ribosomal subunit (30S) binds to the mRNA at the Shine-Dalgarno sequence, a purine-rich sequence (typically AGGAGG) located 6–10 nucleotides upstream of the start codon. This sequence base-pairs with the anti-Shine-Dalgarno sequence at the 3′ end of the 16S ribosomal RNA (rRNA).
  2. The initiator tRNA, charged with formylmethionine (fMet-tRNAfMet), binds to the 30S subunit in the P (peptidyl) site, guided by initiation factor 2 (IF2) bound to GTP.
  3. The 30S initiation complex scans or directly positions the start codon in the P site. The codon–anticodon interaction between AUG and the anticodon of fMet-tRNAfMet (3′-UAC-5′) stabilizes the complex.
  4. Initiation factor 3 (IF3) ensures that the initiator tRNA binds only to the P site and only in response to a start codon, preventing premature binding of elongator tRNAs.
  5. The 50S large subunit joins, GTP is hydrolyzed, and initiation factors are released, forming the 70S initiation complex. The ribosome is now ready for the elongation phase.

In eukaryotes, the process is more elaborate:

  1. The small ribosomal subunit (40S) binds to the 5′ cap structure of the mRNA, along with initiation factors eIF1, eIF1A, eIF3, and the eIF2-GTP-Met-tRNAi ternary complex.
  2. The 43S preinitiation complex scans along the 5′ untranslated region (UTR) in a 5′ to 3′ direction, unwinding secondary structures with the help of eIF4A, an RNA helicase.
  3. The complex recognizes the first AUG in a favorable context as the start codon. Recognition involves base-pairing between the AUG and the anticodon of Met-tRNAi, as well as inspection of the surrounding nucleotide context by eIF1 and eIF1A.
  4. Upon start codon recognition, eIF2-bound GTP is hydrolyzed, eIF2-GDP is released, and the 60S subunit joins to form the 80S ribosome.

Kozak Sequence in Eukaryotes

In eukaryotes, the nucleotide context surrounding the start codon strongly influences initiation efficiency. The Kozak consensus sequence, named after Marilyn Kozak who characterized it in the 1980s, is GCCRCCAUGG (where R is a purine, A or G). The most critical positions are the purine at position −3 (three nucleotides upstream of the AUG) and the G at position +4 (immediately downstream of the AUG). A strong Kozak context (GCCACCAUGG) supports efficient initiation, while a weak context (e.g., UUUAUGAUC) reduces initiation efficiency by 5- to 10-fold.

The scanning mechanism means that the first AUG encountered by the 40S subunit is usually the start codon. However, if the first AUG is in a weak context, some scanning ribosomes may bypass it and initiate at a downstream AUG in a stronger context—a phenomenon called leaky scanning. This provides a mechanism for producing multiple protein isoforms from a single mRNA.

Shine-Dalgarno Sequence in Prokaryotes

In prokaryotes, the Shine-Dalgarno (SD) sequence is the primary determinant of start codon selection. The consensus SD sequence is AGGAGG, located 5–9 nucleotides upstream of the start codon. It base-pairs with the anti-SD sequence (CCUCCU) at the 3′ end of the 16S rRNA. The spacing between the SD sequence and the start codon is critical: optimal spacing is 5–9 nucleotides, and deviations reduce initiation efficiency substantially.

The SD sequence positions the start codon directly in the P site of the 30S subunit, eliminating the need for scanning. This is why prokaryotic mRNAs are typically polycistronic—they contain multiple genes, each with its own SD sequence and start codon, allowing independent translation of each coding sequence from a single mRNA transcript.

The Initiator tRNA and Methionine

The initiator tRNA is a specialized tRNA that is dedicated to translation initiation. In bacteria, it is designated tRNAfMet and carries a modified methionine called formylmethionine (fMet). The formyl group is added to the amino group of methionine by the enzyme methionyl-tRNA formyltransferase, using N10-formyltetrahydrofolate as the formyl donor. The formyl group serves two purposes: it prevents fMet from entering the elongation cycle (since elongator Met-tRNA is not formylated), and it protects the N-terminus of the nascent polypeptide from degradation by aminopeptidases during early elongation.

In eukaryotes, the initiator tRNA is designated tRNAiMet. It carries unmodified methionine, but it has several distinguishing features:

  • It has a unique anticodon loop structure that allows it to bind directly to the P site of the ribosome without prior binding to the A site.
  • It is recognized specifically by eIF2, which delivers it to the 40S subunit as part of the ternary complex.
  • It lacks the base-pair between positions 1 and 72 that is present in elongator tRNAs, giving it a distinctive cloverleaf structure.
  • It is not recognized by elongation factor 1A (eEF1A), ensuring that it does not participate in elongation.

The initiator tRNA is charged with methionine by methionyl-tRNA synthetase, the same enzyme that charges elongator tRNAMet. The distinction between initiator and elongator tRNAMet is made by the ribosome and initiation factors, not by the amino acid itself. This is why alternative start codons like GUG and UUG still result in methionine at the N-terminus: the initiator tRNA's anticodon (3′-UAC-5′) can wobble-pair with these codons, and the initiation machinery tolerates this non-canonical pairing.

After initiation, the N-terminal methionine is often removed co-translationally by methionine aminopeptidase, and in bacteria, the formyl group is removed by deformylase. Approximately 50–70% of mature proteins in E. coli lack the N-terminal methionine, and a similar proportion in eukaryotes undergo N-terminal methionine excision.

Experimental Methods to Identify Start Codons

Identifying the precise start codon of a gene is essential for understanding protein function, predicting protein sequences, and designing expression constructs. Several experimental approaches are used.

Ribosome Profiling

Ribosome profiling (also called Ribo-seq) is a high-throughput technique that captures the positions of ribosomes on mRNAs genome-wide. The method involves:

  1. Treating cells with cycloheximide (100 µg/mL) to freeze ribosomes on mRNAs.
  2. Digesting unprotected mRNA with RNase I, leaving only the ~28–30 nucleotide fragments protected by the ribosome.
  3. Purifying these ribosome-protected fragments (RPFs), converting them to a cDNA library, and deep-sequencing them.
  4. Aligning the sequenced fragments to the genome to determine ribosome occupancy at nucleotide resolution.

The start codon appears as a strong peak of ribosome occupancy at the 5′ end of the coding sequence, with a characteristic triplet periodicity that reflects the reading frame. Ribosome profiling can identify both annotated and unannotated start codons, including those at non-AUG codons, and can reveal translation of upstream open reading frames (uORFs) in 5′ UTRs.

Reporter Constructs

Reporter gene assays are a classical approach for testing whether a candidate sequence can initiate translation. The strategy involves:

  1. Cloning the candidate 5′ UTR and the first few codons of the gene of interest upstream of a reporter gene such as luciferase, GFP (green fluorescent protein), or lacZ (β-galactosidase).
  2. Transfecting or transforming the construct into appropriate cells.
  3. Measuring reporter activity. If the candidate sequence contains a functional start codon, the reporter will be translated and produce a measurable signal.

For example, to test whether a CUG codon can serve as a start codon, one would construct a plasmid with the CUG in frame with the reporter gene, with no upstream AUG codons. If reporter activity is detected, the CUG is functional as a start codon. This approach can be made quantitative by using a dual-luciferase system, where the test construct is fused to firefly luciferase and a control construct (with a known strong start codon) is fused to Renilla luciferase. The ratio of firefly to Renilla activity normalizes for transfection efficiency and cell viability.

N-terminal sequencing, using Edman degradation or mass spectrometry, provides direct evidence of the start codon by determining the first few amino acids of the mature protein. Comparison of the experimentally determined N-terminal sequence with the predicted protein sequence from genomic DNA reveals the exact start codon used. This method is low-throughput but definitive.

Common Misconceptions and Pitfalls

Several misconceptions about start codons are common among students and can lead to errors in exam answers and experimental design.

Misconception 1: All start codons are AUG. While AUG is the most common and the canonical start codon, alternative start codons (GUG, UUG, CUG, AUA, AUU) exist in various organisms and organelles. In E. coli, approximately 8% of genes use GUG and 1% use UUG. The key point is that regardless of the start codon sequence, the initiating amino acid is methionine (or formylmethionine in bacteria).

Misconception 2: The start codon is the same as the promoter. The promoter is a DNA sequence upstream of the gene that recruits RNA polymerase for transcription. The start codon is an RNA sequence (AUG) within the mRNA that recruits the ribosome for translation. These are fundamentally different signals operating at different stages of gene expression. A promoter is never translated; a start codon is never transcribed as DNA (though it is encoded in DNA as ATG).

Misconception 3: Translation reads the mRNA from the 3′ end. Translation always proceeds in the 5′ to 3′ direction along the mRNA. The start codon is the first codon at the 5′ end of the coding region, and the ribosome moves toward the 3′ end, encountering Termination Codon signals at the end of the coding sequence. This directionality is a consequence of the structure of the ribosome and the chemistry of peptide bond formation.

Misconception 4: The start codon codes for methionine in all contexts. In the standard genetic code, AUG codes for methionine. However, in mitochondria with variant genetic codes, AUG may not always specify methionine, and in some mitochondrial systems, AUA codes for methionine rather than isoleucine. Additionally, when AUG is used as a start codon, the initiating methionine is often removed post-translationally, so the mature protein may not begin with methionine.

Misconception 5: The first AUG in the mRNA is always the start codon. In eukaryotes, the first AUG is usually the start codon, but if it is in a weak Kozak context, scanning ribosomes may bypass it (leaky scanning) and initiate at a downstream AUG. In prokaryotes, the start codon is defined by its proximity to the Shine-Dalgarno sequence, not by its position in the mRNA. An AUG without a properly spaced SD sequence will not initiate translation efficiently.

Pitfall in experimental design: Ignoring codon context. When designing expression constructs, placing a start codon in a poor context can reduce protein expression by 5- to 10-fold. For optimal expression in mammalian cells, the sequence GCCACCATGG (with the start codon underlined) should be used. For E. coli, ensure that the SD sequence and start codon are separated by 5–9 nucleotides.

Pitfall in bioinformatics: Assuming the first ATG in genomic DNA is the start codon. Genomic DNA contains introns in eukaryotes, and the first ATG in the genomic sequence may be in an intron or in the 5′ UTR. Start codon prediction requires consideration of splice sites, Kozak context, and comparison with homologous genes. Tools that incorporate these features, such as ORFfinder and GeneMark, are more reliable than simple ATG-scanning algorithms.

Frequently Asked Questions

What is the starting codon?

The starting codon is the first codon of an mRNA that is translated by the ribosome. In the standard genetic code, it is AUG, which codes for methionine. The starting codon establishes the reading frame for the entire protein-coding sequence.

Why is AUG the start codon?

AUG is the start codon because the initiator tRNA, which carries methionine, has an anticodon (UAC) that is perfectly complementary to AUG. This provides the strongest codon–anticodon interaction during initiation, ensuring accurate start site selection. The universality of AUG across nearly all organisms reflects its evolutionary fixation early in the history of life.

Are there other start codons besides AUG?

Yes. Alternative start codons include GUG, UUG, and CUG in bacteria and some eukaryotes, and AUA and AUU in mitochondria. These codons are used less efficiently than AUG, and when they serve as start codons, they still direct the incorporation of methionine (or formylmethionine in bacteria) because the initiator tRNA can wobble-pair with them.

What is the difference between start codon and promoter?

The promoter is a DNA sequence located upstream of a gene that binds RNA polymerase to initiate transcription. The start codon is an RNA sequence (AUG) within the mRNA that signals the ribosome to begin translation. The promoter controls whether a gene is transcribed; the start codon controls where translation begins on the resulting mRNA.

Does the start codon always code for methionine?

In the standard genetic code, yes—AUG always codes for methionine. However, the initiating methionine is often removed from the mature protein by methionine aminopeptidase. In bacteria, the initiating amino acid is formylmethionine, which is also usually removed. In variant genetic codes, such as those in some mitochondria, AUG may not code for methionine in all contexts.

How is the start codon recognized by the ribosome?

In prokaryotes, the 30S ribosomal subunit binds to the Shine-Dalgarno sequence upstream of the start codon, positioning the AUG in the P site where it base-pairs with the anticodon of fMet-tRNAfMet. In eukaryotes, the 40S subunit scans from the 5′ cap of the mRNA until it encounters an AUG in a favorable Kozak context, where it base-pairs with the anticodon of Met-tRNAi.

What happens if the start codon is mutated?

If the start codon is mutated, translation initiation is impaired or abolished. A mutation that changes AUG to a non-start codon (e.g., AUG to AUA) will prevent ribosome binding at that site, and translation may initiate at a downstream AUG, producing a truncated protein missing its N-terminal portion. If no downstream AUG is available, no protein is produced. Such mutations are often deleterious and can cause genetic diseases. For example, mutations in the start codon of the HBB gene (encoding β-globin) cause β-thalassemia due to reduced or absent β-globin synthesis.

Key Takeaways

  • The starting codon is AUG in the standard genetic code, and it codes for methionine while also defining the reading frame for the entire protein.
  • Alternative start codons (GUG, UUG, CUG) exist but are used less efficiently and still result in methionine at the N-terminus.
  • Start codon recognition differs between prokaryotes (Shine-Dalgarno sequence) and eukaryotes (Kozak sequence and scanning mechanism).
  • The initiator tRNA is a specialized molecule with unique features that restrict it to initiation, not elongation.
  • The start codon is an RNA signal for translation; the promoter is a DNA signal for transcription—they are distinct and should not be confused.
  • Mutations in the start codon typically abolish or reduce protein synthesis and can cause disease.
  • Experimental identification of start codons relies on ribosome profiling, reporter assays, and N-terminal sequencing, each with specific strengths and limitations.

Related Clinical & Scientific Guides