Start Codon: Definition, Role in Translation, and Examples

By Dr. Zubair Khalid, DVM, MS, PhD ·

Start Codon: Definition, Role in Translation, and Examples

Introduction to the Start Codon

Definition and Basic Function

A start codon is the first codon of a messenger RNA (mRNA) transcript that is translated into protein. It serves two essential functions: it signals the ribosome to begin protein synthesis, and it establishes the reading frame—the grouping of nucleotides into consecutive triplets—that the ribosome will maintain for the remainder of the transcript. Because the genetic code is read in non-overlapping groups of three nucleotides, the position of the start codon determines which of the three possible reading frames is used. A shift of even one nucleotide at this stage would produce a completely different, and almost certainly non-functional, protein.

The start codon is recognized by the translation machinery during the initiation phase of protein synthesis, which is the most highly regulated step of translation. In both prokaryotes and eukaryotes, initiation requires the coordinated assembly of the small ribosomal subunit, the initiator transfer RNA (tRNA), and various initiation factors at the correct position on the mRNA. The start codon is the anchor point for this assembly.

The Standard Start Codon: AUG

The canonical start codon is AUG, which codes for the amino acid methionine. This codon is nearly universal across all domains of life. In bacteria, the initiating amino acid is a modified form of methionine called N-formylmethionine (fMet), which is attached to a specialized initiator tRNA. In eukaryotes and archaea, the initiating amino acid is unmodified methionine, though it is still carried by a dedicated initiator tRNA that is distinct from the tRNA used to insert methionine at internal positions in a polypeptide chain.

AUG is not exclusively a start codon; it also encodes methionine when it appears in the interior of an open reading frame (ORF). This dual role means that the translation machinery must distinguish between AUG codons that serve as initiation sites and those that encode internal methionine residues. This discrimination is achieved through sequence context, as discussed in later sections.

The Genetic Code and Start Codons

Codon Table and Start Codons

The genetic code is a set of 64 possible codons—three-nucleotide sequences—that specify 20 amino acids and three stop signals. Of these 64 codons, 61 are sense codons that specify amino acids, and three are Stop Codon (UAA, UAG, UGA) that terminate translation. AUG is the sole codon that functions as the standard start signal in the vast majority of organisms. The Codon Table lists AUG under methionine, and it is the only codon with a dual role as both an initiation signal and an internal amino acid codon.

The code is degenerate, meaning that most amino acids are specified by more than one codon. Methionine, however, is specified only by AUG. This uniqueness simplifies the identification of start codons in genomic sequences: a long ORF that begins with AUG and ends with a stop codon is a strong candidate for a protein-coding gene. The Codon Wheel provides a visual representation of this relationship, showing AUG at the center of the methionine position.

Methionine as the Initiating Amino Acid

The choice of methionine as the initiating amino acid is not arbitrary. Methionine has a hydrophobic side chain that fits into the ribosomal P site during initiation, and its sulfur atom plays a role in the stability of the initiator tRNA–ribosome interaction. In bacteria, the formyl group added to the amino group of methionine prevents the initiating amino acid from being recognized by elongation factors, ensuring that only the initiator tRNA can enter the P site during initiation.

After translation begins, the N-terminal methionine is often removed by specific aminopeptidases. In eukaryotes, approximately 60–70% of proteins undergo this cleavage, a process called N-terminal methionine excision. In bacteria, the formyl group is removed by deformylase, and the methionine itself may then be removed by methionine aminopeptidase. This post-translational processing means that the final protein product frequently does not retain the initiating methionine, which can complicate the identification of start codons from protein sequence data alone.

Mechanism of Translation Initiation

Initiation Complex Formation

Translation initiation proceeds through a series of ordered steps that culminate in the assembly of a ribosome positioned precisely at the start codon. The process differs between prokaryotes and eukaryotes in its details, but the fundamental logic is conserved.

In prokaryotes, initiation involves the 30S ribosomal subunit, three initiation factors (IF1, IF2, and IF3), mRNA, and the initiator tRNA carrying fMet. The steps are as follows:

  1. IF3 binds to the 30S subunit, preventing premature association with the 50S subunit and promoting the binding of mRNA.
  2. IF1 binds near the A site of the 30S subunit, blocking the A site and enhancing the activity of IF3.
  3. The mRNA binds to the 30S subunit through the Shine-Dalgarno sequence, which base-pairs with the anti-Shine-Dalgarno sequence at the 3′ end of 16S ribosomal RNA.
  4. IF2, bound to GTP, delivers the initiator tRNA (tRNAfMet) to the P site of the 30S subunit, where it base-pairs with the AUG start codon.
  5. IF3 is released, and the 50S subunit joins the complex, triggering GTP hydrolysis by IF2 and the release of IF1 and IF2.
  6. The complete 70S ribosome is now assembled with the initiator tRNA in the P site, positioned over the start codon, and the A site is empty, ready for the next codon.

In eukaryotes, the process is more complex and involves at least 12 initiation factors (eIFs). The key steps are:

  1. The 40S ribosomal subunit binds eIF1, eIF1A, eIF3, and the eIF2–GTP–Met-tRNAi ternary complex, forming the 43S preinitiation complex.
  2. The 43S complex binds the 5′ cap of the mRNA through eIF4F, which includes eIF4E (cap-binding protein), eIF4A (RNA helicase), and eIF4G (scaffold protein).
  3. The complex scans along the mRNA in the 5′ to 3′ direction, unwinding secondary structure with the help of eIF4A.
  4. When the complex encounters an AUG codon in a favorable context, the anticodon of Met-tRNAi base-pairs with the start codon, and eIF1 is released.
  5. This triggers GTP hydrolysis by eIF2, which is a key regulatory checkpoint. The hydrolysis of GTP to GDP causes eIF2 to release from the complex.
  6. eIF5B (the eukaryotic homolog of IF2) promotes the joining of the 60S subunit, and all remaining initiation factors are released.
  7. The complete 80S ribosome is assembled, with Met-tRNAi in the P site, ready for elongation.

Kozak Sequence in Eukaryotes

In eukaryotes, the efficiency of start codon recognition depends heavily on the surrounding nucleotide sequence. The consensus sequence, first characterized by Marilyn Kozak in the 1980s, is GCCRCCaugG, where R is a purine (A or G) at position −3 relative to the A of the start codon, and G at position +4. The most critical positions are the purine at −3 and the G at +4. A strong Kozak sequence (GCCACCaugG) promotes efficient initiation, while a weak context (e.g., a pyrimidine at −3) reduces initiation efficiency by 5- to 10-fold.

The Kozak sequence is recognized by the scanning 40S subunit, and it influences the probability that the ribosome will pause at a given AUG and commit to initiation. If the first AUG encountered is in a weak context, the scanning ribosome may bypass it and initiate at a downstream AUG—a phenomenon called leaky scanning. This provides a mechanism for producing multiple protein isoforms from a single mRNA.

Shine-Dalgarno Sequence in Prokaryotes

In prokaryotes, the start codon is recognized through a different mechanism. The Shine-Dalgarno (SD) sequence, with the consensus AGGAGG, is located 5–10 nucleotides upstream of the start codon. It base-pairs with the anti-Shine-Dalgarno sequence (CCUCCU) at the 3′ end of 16S ribosomal RNA. This interaction positions the 30S subunit so that the start codon is placed directly in the P site.

The spacing between the SD sequence and the start codon is critical. A spacing of 5–9 nucleotides is optimal; deviations of even one nucleotide can reduce translation efficiency by 50% or more. The strength of the SD interaction also matters: a stronger SD sequence (more complementary to the anti-SD) generally leads to higher translation initiation rates, but excessively strong interactions can stall the ribosome during initiation.

Alternative Start Codons

Examples of Alternative Start Codons

Although AUG is the standard start codon, several alternative start codons exist in nature. The most common are GUG (valine) and UUG (leucine), which are used as start codons in approximately 1–3% of bacterial genes. In Escherichia coli, for example, about 83% of genes start with AUG, 14% with GUG, and 3% with UUG. Other rare alternatives include CUG (leucine), AUU (isoleucine), and AUA (isoleucine).

When a non-AUG codon serves as the start codon, the initiating amino acid is still methionine, not the amino acid that the codon normally specifies. This is because the initiator tRNA, which carries methionine, can base-pair with these alternative codons with reduced efficiency. The wobble position of the anticodon allows the initiator tRNA to recognize GUG and UUG, albeit less efficiently than AUG. This reduced efficiency means that genes with alternative start codons are typically translated at lower levels than those with AUG.

Alternative start codons are particularly common in viruses and in genes encoding regulatory proteins, where low translation efficiency is functionally important. For example, the λ phage cI repressor gene uses both AUG and GUG start codons to produce two protein isoforms with different functions.

Context-Dependent Recognition

The recognition of alternative start codons is highly context-dependent. A GUG or UUG codon in a strong Shine-Dalgarno context in bacteria, or a strong Kozak context in eukaryotes, can function as an efficient start site. Conversely, an AUG codon in a poor context may be skipped entirely.

In eukaryotes, non-AUG start codons are less common but do occur. CUG is the most frequently used alternative start codon in mammalian cells, and it has been found in several oncogenes and growth factors. For instance, the human c-myc gene can initiate translation at a CUG codon, producing a longer protein isoform with different regulatory properties. The efficiency of CUG initiation is typically 10–20% of that of AUG, but this can be modulated by the Kozak context and by the secondary structure of the mRNA.

Start Codon Recognition and Reading Frame

Importance of Reading Frame

The start codon establishes the reading frame, which is the grouping of nucleotides into codons that the ribosome reads sequentially. Because the genetic code is triplet-based and non-overlapping, there are three possible reading frames in any given mRNA sequence. The start codon selects one of these three frames, and the ribosome maintains this frame throughout elongation.

The importance of the reading frame cannot be overstated. A single nucleotide insertion or deletion downstream of the start codon shifts the reading frame, causing the ribosome to read a completely different sequence of codons from that point onward. This almost always leads to the production of a truncated or non-functional protein, because the shifted frame typically encounters a stop codon prematurely.

Consequences of Frameshift Mutations

Frameshift mutations—insertions or deletions of nucleotides that are not multiples of three—are among the most deleterious types of mutations. Consider a coding sequence that is 300 codons long. A single nucleotide deletion at codon 50 will cause all codons downstream to be read in a shifted frame. The probability that this shifted frame contains a stop codon within the next 50 codons is high, given that stop codons occur roughly once every 20 codons in random sequence. The result is a truncated protein that lacks the C-terminal domains essential for its function.

Frameshift mutations can also occur at the level of translation initiation if the start codon is not recognized correctly. If the ribosome initiates at the wrong AUG, or at a non-AUG codon in the wrong context, the reading frame will be shifted, and the resulting protein will be non-functional. This is why the fidelity of start codon recognition is so critical: errors at this step are catastrophic for protein function.

Methods to Study Start Codons

Ribosome Profiling

Ribosome profiling, also known as Ribo-seq, is a powerful technique for identifying start codons genome-wide. The method involves the following steps:

  1. Cells are treated with cycloheximide (in eukaryotes) or chloramphenicol (in prokaryotes) to halt translation and freeze ribosomes on mRNAs.
  2. The mRNA is digested with nuclease, leaving only the ribosome-protected fragments (RPFs), which are typically 28–30 nucleotides long in eukaryotes and 20–25 nucleotides in prokaryotes.
  3. The RPFs are purified, converted to a cDNA library, and subjected to high-throughput sequencing.
  4. The sequenced fragments are aligned to the genome, and the positions of the ribosome P sites are inferred from the fragment lengths and alignment offsets.

The P site position corresponds to the codon being translated at the moment of freezing. At initiation, the P site is occupied by the start codon. Therefore, a sharp peak of ribosome occupancy at the beginning of an ORF, with a characteristic triplet periodicity, identifies the start codon. Ribosome profiling has revealed that many genes have multiple translation start sites, and it has identified thousands of upstream open reading frames (uORFs) that begin with non-AUG codons.

Site-Directed Mutagenesis

Site-directed mutagenesis is a classic approach for studying start codons. The technique involves introducing specific mutations into a cloned gene and measuring the effect on protein expression. For example, to test whether a putative start codon is functional, the researcher can:

  1. Mutate the candidate AUG to a stop codon (e.g., UAA) and observe whether protein production ceases.
  2. Mutate the candidate AUG to a non-AUG codon (e.g., AUA) and measure the reduction in translation efficiency.
  3. Introduce mutations in the Kozak or Shine-Dalgarno sequence to assess the importance of context.

These experiments are typically performed using a reporter gene assay, where the gene of interest is fused to a reporter such as green fluorescent protein (GFP) or luciferase, allowing quantitative measurement of translation efficiency.

Reporter Gene Assays

Reporter gene assays are widely used to study start codon function and context. The basic design is as follows:

  1. The sequence surrounding the putative start codon is cloned upstream of a reporter gene that lacks its own start codon.
  2. The construct is introduced into cells, and reporter activity is measured.
  3. By systematically mutating the start codon and its surrounding sequence, the researcher can determine the efficiency of initiation under different conditions.

For example, to test the strength of a Kozak sequence, a researcher might create a series of constructs with different nucleotides at the −3 and +4 positions and measure luciferase activity. A strong Kozak context (GCCACCaugG) typically yields 5- to 10-fold higher reporter activity than a weak context (e.g., TCTACTaugC). These assays can also be used to test the effects of alternative start codons, RNA secondary structure, and trans-acting factors on initiation efficiency.

Common Misconceptions and Pitfalls

Start Codon vs. Promoter

A frequent source of confusion is the distinction between a start codon and a promoter. The promoter is a DNA sequence located upstream of a gene that serves as the binding site for RNA polymerase and transcription factors. It controls when and how often a gene is transcribed into mRNA. The start codon, by contrast, is an RNA sequence within the mRNA that controls where translation begins. The promoter is not transcribed into the mRNA (or is only partially transcribed in the case of the 5′ untranslated region), whereas the start codon is the first codon of the coding sequence.

Students often conflate these two elements because both are involved in gene expression and both are located near the beginning of a gene. However, they operate at fundamentally different levels: the promoter is DNA and controls transcription; the start codon is RNA and controls translation. Mutations in a promoter affect mRNA levels, while mutations in a start codon affect translation efficiency or completely abolish protein production.

Not All AUGs Are Start Codons

Another common error is assuming that every AUG in an mRNA is a start codon. In a typical human mRNA, the 5′ untranslated region (5′ UTR) may contain one or more AUG codons that are not used as start sites. These upstream AUGs (uAUGs) can regulate translation by causing the ribosome to initiate at an upstream ORF, which is usually short and non-functional, thereby reducing translation of the main ORF.

Furthermore, AUG codons within the coding sequence specify internal methionine residues and are not start codons. The ribosome does not re-initiate at these internal AUGs during normal elongation. The distinction between an initiation AUG and an elongation AUG is determined by the context: the Kozak sequence in eukaryotes and the Shine-Dalgarno sequence in prokaryotes mark the initiation AUG specifically.

Context Matters

A third pitfall is ignoring the importance of sequence context. Students sometimes memorize that AUG is the start codon and assume that any AUG will work equally well. In reality, the efficiency of initiation at an AUG codon depends heavily on its surrounding sequence. An AUG in a poor Kozak context may be skipped by the scanning ribosome, leading to initiation at a downstream AUG. This phenomenon, called leaky scanning, can produce multiple protein isoforms from a single mRNA.

Similarly, in prokaryotes, the Shine-Dalgarno sequence and its spacing from the start codon are critical. A gene with a weak SD sequence or suboptimal spacing will be translated at low efficiency, even if the start codon is AUG. Conversely, a strong SD sequence can compensate for a non-AUG start codon, allowing efficient initiation at GUG or UUG.

Summary and Key Takeaways

The start codon is the foundation of protein synthesis. It marks the beginning of the coding sequence, establishes the reading frame, and is the target of the most highly regulated step in translation. The standard start codon is AUG, which specifies methionine, but alternative start codons such as GUG and UUG are used in a minority of genes. Recognition of the start codon depends on sequence context: the Kozak sequence in eukaryotes and the Shine-Dalgarno sequence in prokaryotes. Errors in start codon selection lead to frameshifts and non-functional proteins, underscoring the importance of this molecular decision point.

Frequently Asked Questions

What is a start codon?

A start codon is the first codon of an mRNA sequence that is translated into protein. It signals the ribosome to begin protein synthesis and establishes the reading frame for the entire coding sequence. The standard start codon is AUG, which codes for methionine.

What is the start codon example?

The most common example of a start codon is AUG, which codes for methionine. In bacteria, the initiating amino acid is N-formylmethionine. Alternative start codons include GUG (normally valine) and UUG (normally leucine), which are used in a small percentage of genes.

What is the meaning of start codon?

The start codon is the specific three-nucleotide sequence in mRNA that marks the position where translation begins. It serves two functions: it recruits the ribosome and initiator tRNA to begin protein synthesis, and it sets the reading frame that determines how all subsequent codons are read.

What are the types of start codons?

The types of start codons include the canonical AUG and several alternative start codons. The most common alternatives are GUG and UUG, which are used in approximately 14% and 3% of E. coli genes, respectively. Rarer alternatives include CUG, AUU, and AUA. In all cases, the initiating amino acid is methionine, regardless of the codon sequence.

How does the start codon process work?

The start codon process involves the assembly of the translation initiation complex. In eukaryotes, the 40S ribosomal subunit scans from the 5′ cap of the mRNA until it encounters an AUG in a favorable Kozak context. The initiator tRNA then base-pairs with the start codon, GTP is hydrolyzed, and the 60S subunit joins to form the complete ribosome. In prokaryotes, the 30S subunit binds directly to the Shine-Dalgarno sequence upstream of the start codon, positioning the AUG in the P site.

Can you show a start codon diagram?

A start codon diagram typically shows the mRNA sequence with the 5′ untranslated region, the start codon (AUG) in the center, and the downstream coding sequence. In eukaryotes, the Kozak consensus sequence (GCCRCCaugG) is shown flanking the start codon. In prokaryotes, the Shine-Dalgarno sequence (AGGAGG) is shown 5–10 nucleotides upstream. The initiator tRNA is depicted base-paired with the start codon in the P site of the ribosome.

What is the list of start codons?

The list of start codons includes:

  • AUG (canonical; methionine)
  • GUG (valine; used as start codon in ~14% of E. coli genes)
  • UUG (leucine; used as start codon in ~3% of E. coli genes)
  • CUG (leucine; used in some eukaryotic genes)
  • AUU (isoleucine; rare)
  • AUA (isoleucine; rare)

In all cases, the initiator tRNA carries methionine, so the first amino acid is always methionine regardless of the start codon sequence.

Key Takeaways

  • The start codon is the first codon of an mRNA coding sequence and is almost always AUG, which specifies methionine.
  • The start codon establishes the reading frame; a shift of one nucleotide at this point produces a completely different protein.
  • Recognition of the start codon depends on sequence context: the Kozak sequence in eukaryotes and the Shine-Dalgarno sequence in prokaryotes.
  • Alternative start codons (GUG, UUG, CUG) exist and are used in a minority of genes, often to regulate translation efficiency.
  • The initiating amino acid is always methionine, even when the start codon is not AUG.
  • Not every AUG in an mRNA is a start codon; internal AUGs encode methionine, and upstream AUGs in the 5′ UTR can regulate translation.
  • Errors in start codon recognition cause frameshift mutations and typically produce non-functional proteins, highlighting the critical role of this molecular decision point.

Further Reading

  • Dever TE, Ivanov IP, Hinnebusch AG. Translational regulation by uORFs and start codon selection stringency. Genes & development. 2023. PubMed 37433636
  • Cao X, Slavoff SA. Non-AUG start codons: Expanding and regulating the small and alternative ORFeome. Experimental cell research. 2020. PubMed 32209305
  • Ly J et al. Nuclear release of eIF1 restricts start-codon selection during mitosis. Nature. 2024. PubMed 39443796
  • Mao Y et al. Start codon-associated ribosomal frameshifting mediates nutrient stress adaptation. Nature structural & molecular biology. 2023. PubMed 37957305
  • Ly J et al. Alternative start codon selection shapes mitochondrial function and rare human diseases. Molecular cell. 2025. PubMed 41205602
  • Asano K. Why is start codon selection so precise in eukaryotes?. Translation (Austin, Tex.). 2014. PubMed 26779403

Related Clinical & Scientific Guides