E. coli Codon Optimization: Mechanisms and Best Practices
By Dr. Zubair Khalid, DVM, MS, PhD ·

Introduction to Codon Optimization in E. coli
The Genetic Code and Codon Degeneracy
The genetic code translates the four-letter nucleotide alphabet into the twenty canonical amino acids through triplet codons. With 64 possible codons and only 20 amino acids plus a translation stop signal, the code is necessarily degenerate: most amino acids are specified by multiple synonymous codons. Leucine, arginine, and serine are each encoded by six codons; isoleucine by three; and methionine and tryptophan by exactly one. This degeneracy is not random noise in the code—it is a substrate for evolutionary tuning of gene expression.
Synonymous codons are not functionally equivalent. They differ in their interaction with the anticodon of cognate tRNAs, their base composition, their position-dependent effects on mRNA secondary structure, and their influence on translation kinetics. The choice among synonymous codons is therefore under selective pressure, and organisms display characteristic codon usage biases that reflect their genomic history, tRNA repertoire, and growth environment.
Why Optimize for E. coli?
Escherichia coli is the most widely used prokaryotic host for recombinant protein production. Its rapid doubling time, well-characterized genetics, inexpensive culture requirements, and extensive toolkit of expression plasmids make it the default first choice for heterologous expression. However, genes from other organisms—particularly eukaryotes—often express poorly in E. coli when their coding sequences are transferred directly.
The primary reason is codon usage mismatch. A gene evolved under the codon bias of Saccharomyces cerevisiae, Homo sapiens, or a thermophilic archaeon will contain codons that are rare in E. coli. When the translation machinery encounters these rare codons, elongation slows or stalls, leading to reduced yield, truncated products, and protein misfolding. Codon optimization addresses this by redesigning the coding sequence to match the host's translational preferences while preserving the amino acid sequence.
Optimization also addresses secondary issues: mRNA secondary structure near the ribosome binding site, GC content extremes that affect transcript stability, and the inadvertent introduction of regulatory motifs that destabilize mRNA or prematurely terminate transcription.
Codon Usage Bias in E. coli
tRNA Pool and Codon Availability
E. coli K-12 carries approximately 86 tRNA genes distributed across the chromosome. The abundance of each tRNA species in the cell is not uniform; tRNAs for codons used frequently in highly expressed genes are present at higher concentrations. This correlation is not coincidental—it reflects co-evolution of the genome and the translation apparatus under selection for efficient protein production.
The relationship between codon usage and tRNA abundance is quantified by the tRNA Adaptation Index (tAI), which weights each codon by the copy number and anticodon pairing efficiency of its cognate tRNA. For E. coli, the most abundant tRNAs correspond to codons such as CTG (Leu), GGT (Gly), and GAA (Glu), while tRNAs for codons like AGG/AGA (Arg), ATA (Ile), and CTA (Leu) are scarce.
The pool of charged tRNAs is dynamic. Under amino acid limitation, the fraction of charged tRNA drops, and codons requiring scarce tRNAs become particularly problematic. This is why expression of AT-rich genes from organisms like Plasmodium falciparum is especially challenging in E. coli—the high A/T content creates a preponderance of codons that are both rare and dependent on limiting tRNA species.
Measuring Codon Usage Bias: CAI and Other Indices
The Codon Adaptation Index (CAI) is the most widely used metric for quantifying how well a coding sequence matches a reference set of highly expressed genes. CAI is calculated as the geometric mean of relative synonymous codon usage (RSCU) values, where each codon's RSCU is its observed frequency divided by the frequency expected if all synonymous codons for that amino acid were used equally. A CAI of 1.0 indicates perfect match to the reference set; values above 0.8 are generally considered good for E. coli expression.
Other indices include the Effective Number of Codons (Nc), which measures codon usage uniformity without a reference set, and the Codon Pair Bias, which quantifies the non-random frequency of adjacent codon pairs. The latter reflects the observation that some codon pairs are over- or under-represented in highly expressed genes, likely due to effects on translation speed at the dipeptide level and on mRNA structure.
For practical purposes, CAI remains the standard benchmark. Most optimization tools report the CAI of their output, and a CAI above 0.9 is typically achievable for E. coli without compromising other design constraints.
Mechanisms Linking Codon Usage to Protein Expression
Translation Elongation Dynamics
Translation elongation in E. coli proceeds at roughly 15–20 amino acids per second at 37°C, but this rate is not uniform along the mRNA. Codons with abundant cognate tRNAs are decoded rapidly; rare codons cause the ribosome to pause while awaiting a matching aminoacyl-tRNA. These pauses are not merely delays—they have functional consequences.
A stalled ribosome can trigger the transfer-messenger RNA (tmRNA) system, also known as trans-translation. The tmRNA–SmpB complex recognizes stalled ribosomes, adds a proteolytic tag to the nascent polypeptide, and releases the ribosome for recycling. The tagged protein is then degraded by ClpXP or other proteases. Thus, clusters of rare codons can lead to both reduced yield and the appearance of truncated degradation products on SDS-PAGE.
Pausing also affects co-translational processes. The ribosome-associated chaperone trigger factor binds nascent chains as they emerge from the exit tunnel. A prolonged pause allows the nascent chain more time to fold or to interact with chaperones before the next domain emerges. In some cases, this is beneficial; in others, it promotes misfolding or aggregation.
Codon Optimality and mRNA Degradation
The stability of mRNA in E. coli is governed by a balance between ribosome loading and RNase activity. The major mRNA decay pathway begins with RNase E, an endoribonuclease that cleaves single-stranded AU-rich regions. Ribosomes protect mRNA from RNase E cleavage; when ribosomes are sparse or stalled, the mRNA becomes more accessible to degradation.
Codon optimality influences this process directly. The RNA degradosome, a multi-enzyme complex containing RNase E, polynucleotide phosphorylase, and the DEAD-box helicase RhlB, preferentially targets transcripts with non-optimal codons. This is because non-optimal codons slow ribosome transit, creating gaps in ribosome coverage that expose cleavage sites. The result is a positive feedback loop: non-optimal codons reduce translation, which reduces mRNA stability, which further reduces protein output.
This mechanism explains why codon optimization can improve expression even when the original gene has no problematic rare codon clusters. The overall codon optimality of the transcript sets a baseline for its half-life, and optimization raises that baseline.
Co-translational Folding and Protein Aggregation
Protein folding in E. coli begins co-translationally. The ribosome exit tunnel accommodates approximately 30–40 amino acids, and domains can begin folding as soon as they emerge. The rate of translation therefore influences folding outcomes: a fast-translating sequence may produce a domain that folds before a downstream interaction partner is available, while a slow-translating sequence may allow premature folding of an intermediate that is kinetically trapped.
Rare codons are often positioned at domain boundaries in native genes, creating programmed pauses that allow the nascent chain to fold one domain at a time. When such genes are codon-optimized without regard to these pause sites, the resulting protein may misfold and aggregate into inclusion bodies. This is a well-documented failure mode of aggressive optimization.
Conversely, the introduction of pauses at inappropriate positions—for example, within a domain that requires rapid, cooperative folding—can also cause misfolding. The optimal strategy is not to eliminate all pauses but to reposition them to match the folding landscape of the protein.
Codon Optimization Strategies
Full Optimization vs. 'Humanized' Codons
Full optimization replaces every codon with the most frequently used synonymous codon in E. coli for that amino acid. This maximizes CAI and typically produces the highest translation initiation rates. However, it also creates a monotonous sequence with highly repetitive codon pairs, which can have unintended consequences.
The most frequent E. coli codons are AT-rich in the third position for some amino acids and GC-rich for others. A fully optimized sequence may have extreme GC content, particularly in the third codon position, which can stabilize mRNA secondary structures that impede ribosome binding or processivity. Additionally, runs of identical codons can cause ribosome frameshifting, particularly at runs of AGG/AGA (Arg) or CCC (Pro).
A more balanced approach, sometimes called "humanized" or "harmonized" optimization, adjusts the codon usage profile to match the overall distribution of the host genome rather than just the most abundant codons. This produces a sequence with a CAI of 0.85–0.95, moderate GC content, and a more natural distribution of codon pairs. For many proteins, this approach yields equivalent or better expression than full optimization, with fewer folding problems.
Codon Pair Optimization
Codon pair bias refers to the observation that certain adjacent codon pairs occur more or less frequently than expected from the individual codon frequencies. In E. coli, the codon pair CTG-GAA (Leu-Glu) is over-represented, while CTG-GAT (Leu-Asp) is under-represented. The mechanistic basis is not fully resolved but likely involves the kinetics of tRNA release and re-binding at the ribosomal A and P sites.
Codon pair optimization aims to match the codon pair frequency distribution of highly expressed E. coli genes. This is more computationally demanding than single-codon optimization because the design space is larger, but it can improve expression for proteins that are sensitive to translation kinetics. Some tools, such as the Codon Pair Bias tool from the Biotech Research Institute, implement this approach.
The evidence for codon pair effects in E. coli is less robust than in mammalian systems, where codon pair optimization has been used to attenuate viruses. Nevertheless, avoiding extreme codon pair outliers is a reasonable design principle, particularly for proteins expressed at high levels.
Avoiding Rare Codons and Shine-Dalgarno-like Sequences
Rare codons are not uniformly problematic. A single rare codon in an otherwise optimized sequence is unlikely to cause measurable problems. However, clusters of two or more rare codons in close proximity—particularly AGG/AGA (Arg), ATA (Ile), and CTA (Leu)—can cause ribosome stalling and tmRNA tagging. The threshold for problematic clusters is approximately three rare codons within a window of ten codons.
A second motif to avoid is the Shine-Dalgarno (SD) sequence, AGGAGG, or its complement. An SD-like sequence within the coding region can cause the 30S ribosomal subunit to bind internally, leading to translation initiation at an internal AUG and production of N-terminally truncated proteins. The sequence CCUCCU, the complement of the SD sequence, is equally problematic because it can base-pair with the anti-SD sequence of the 16S rRNA.
Optimization algorithms should scan for and eliminate these motifs, along with runs of four or more identical nucleotides, which can cause polymerase slippage during transcription.
Computational Tools for Codon Optimization
Key Parameters in Optimization Algorithms
Modern codon optimization tools balance multiple, sometimes competing, objectives. The key parameters are:
- Codon usage table: The reference set of codon frequencies, typically derived from the E. coli K-12 genome or from a subset of highly expressed genes.
- CAI target: The desired minimum CAI, usually 0.8–0.95.
- GC content: The target GC percentage, typically 40–55% for E. coli.
- mRNA secondary structure: The maximum allowed free energy of folding near the ribosome binding site and the start codon.
- Restriction site avoidance: A list of restriction enzyme recognition sites to eliminate for subsequent cloning.
- Repeat sequence avoidance: A maximum allowed length for direct or inverted repeats.
- Codon pair bias: Whether to optimize for preferred codon pairs or to avoid rare pairs.
The optimization algorithm searches the sequence space of synonymous codons to find a sequence that satisfies these constraints. Most tools use a greedy or Monte Carlo approach, iteratively replacing codons and evaluating the objective function.
Comparing Tool Outputs
Several tools are widely used for E. coli codon optimization:
| Tool | Approach | Key Features | Output Format |
|---|---|---|---|
| IDT Codon Optimization | Multi-objective optimization | GC content, restriction sites, repeats | DNA sequence with codon usage report |
| SnapGene | Single-codon replacement | User-selectable codon table | Sequence with CAI calculation |
| Optimizer | Greedy algorithm | Codon pair optimization option | Sequence with CAI and GC report |
| JCAT (Java Codon Adaptation Tool) | Multi-objective genetic algorithm | CAI, GC, restriction sites, Rho-independent terminators | Sequence with detailed report |
| GenSmart | Machine learning-based | Predicts expression from sequence features | Sequence with expression score |
The outputs of different tools for the same gene can differ substantially. A comparison of JCAT and IDT outputs for a typical eukaryotic gene will show differences in 10–20% of codon positions. Neither is universally superior; the choice depends on the specific constraints of the project. For example, if the gene must be cloned using a specific restriction site, IDT or SnapGene's constraint features are valuable. If the protein is prone to misfolding, a tool that preserves codon pair bias may be preferable.
The Codon Optimization Tool Free resource provides a comparison of freely available options, which is useful for academic laboratories with limited budgets.
Experimental Validation and Case Studies
Expression Systems and Reporter Assays
Codon optimization is only one component of a successful expression strategy. The choice of expression system is equally important. The T7 promoter system, using the pET series of plasmids, is the most common for high-level expression in E. coli. T7 RNA polymerase transcribes several-fold faster than E. coli RNA polymerase, which can exacerbate problems with rare codons because the mRNA is produced faster than the ribosomes can translate it.
For initial screening of optimized genes, a small-scale expression test is recommended. Transform the expression plasmid into E. coli BL21(DE3), grow at 37°C in LB medium to an OD600 of 0.6, induce with 0.5 mM IPTG, and continue growth for 3–4 hours. Analyze whole-cell lysates by SDS-PAGE, comparing the induced and uninduced samples. A soluble fraction should be prepared by sonication in 50 mM Tris-HCl, pH 8.0, 150 mM NaCl, followed by centrifugation at 20,000 × g for 20 minutes at 4°C.
For quantitative comparison of different optimization strategies, a fluorescent reporter is invaluable. Fuse the gene of interest to GFP or use a bicistronic construct with GFP as a translation reporter. Flow cytometry or microplate fluorometry provides a rapid readout of expression levels across many variants.
Case Study: GFP and Therapeutic Proteins
Green fluorescent protein (GFP) from Aequorea victoria is a classic example where codon optimization dramatically improved expression in E. coli. The wild-type gene contains numerous codons rare in E. coli, including AGG (Arg) and ATA (Ile), and expresses poorly. The optimized variant, GFPmut3 or the "cycle3" variant, replaces these with optimal codons and includes mutations that improve folding at 37°C. The result is a several-hundred-fold increase in fluorescence per cell.
Therapeutic proteins present a more complex picture. Human erythropoietin (EPO) is heavily glycosylated in its native context, and E. coli cannot perform this post-translational modification. Codon optimization of EPO for E. coli improves the yield of the polypeptide, but the product is biologically inactive without glycosylation. This illustrates a fundamental limitation: codon optimization cannot compensate for missing post-translational machinery.
Insulin-like growth factor 1 (IGF-1) is a more successful case. This small, non-glycosylated protein expresses well in E. coli after codon optimization and is produced commercially as a fusion protein with a cleavable tag. The optimized gene achieves a CAI above 0.9 and expresses at levels exceeding 100 mg/L in fed-batch fermentation.
Limitations and When Not to Optimize
Rare Codons and Regulatory Roles
Rare codons are not always detrimental. In some genes, they serve regulatory functions. For example, the E. coli gene infA encodes translation initiation factor IF1 and contains a cluster of rare arginine codons (AGG/AGA) that slows its own translation, providing autoregulation. Similarly, the lacZ gene contains a rare proline codon (CCC) that is important for proper folding of the β-galactosidase tetramer.
When expressing a heterologous gene that has evolved under selection for specific translation kinetics, aggressive codon optimization can disrupt these regulatory elements. This is particularly relevant for genes from organisms with similar codon usage to E. coli, such as other enterobacteria. In such cases, the native sequence may already be well-adapted, and optimization provides little benefit while risking disruption of regulatory features.
Membrane Protein Expression Challenges
Membrane proteins are notoriously difficult to express in E. coli, and codon optimization is not a panacea. The bottlenecks for membrane protein expression are often at the level of membrane insertion and folding, not translation. Overexpression of membrane proteins saturates the Sec translocon and the YidC insertase, leading to accumulation of the protein in the cytoplasm as inclusion bodies.
For membrane proteins, a more effective strategy is often to slow translation by using a weaker promoter, lower growth temperature (e.g., 25°C or 30°C), or a strain with reduced tRNA levels. Some studies have found that mild codon optimization—raising the CAI from 0.5 to 0.7 without going to 0.95—improves membrane protein yields by reducing translational pausing without overwhelming the insertion machinery.
Common Pitfalls in E. coli Codon Optimization
Over-Optimization and Codon Pair Bias
The most common mistake is over-optimization. Chasing a CAI of 1.0 often produces a sequence with extreme codon pair bias, high GC content, and strong mRNA secondary structures. The resulting protein may express at lower levels than a moderately optimized sequence, or may misfold into inclusion bodies.
A CAI of 0.85–0.95 is generally sufficient for high expression in E. coli. Beyond this, the marginal gains in translation efficiency are offset by the costs of altered mRNA structure and codon pair distribution. If a sequence with a CAI of 0.9 does not express well, the problem is likely not codon usage but rather mRNA structure, protein toxicity, or a problem with the expression construct.
Ignoring mRNA Secondary Structure
The 5' end of the mRNA, including the ribosome binding site and the first 30–50 codons, must be relatively unstructured for efficient translation initiation. A common failure mode is an optimized sequence that forms a stable stem-loop structure at the 5' end, occluding the Shine-Dalgarno sequence and the start codon.
Most optimization tools include a constraint on mRNA folding energy, but the default settings may not be stringent enough. A free energy of folding (ΔG) of −5 kcal/mol or less in the region from −20 to +30 relative to the start codon is a reasonable target. If the tool does not provide this constraint, check the predicted structure manually using RNAfold or mFold.
Introducing Unintended Motifs
Optimization can inadvertently introduce sequences that are recognized by host regulatory systems. Examples include:
- Rho-independent terminators: GC-rich hairpins followed by a poly-U tract. These can prematurely terminate transcription.
- Shine-Dalgarno-like sequences: Internal AGGAGG motifs that cause spurious translation initiation.
- Restriction sites: Recognition sequences for the restriction enzymes used in cloning, which prevent proper digestion.
- Poly-A tracts: Runs of five or more A residues, which can cause RNA polymerase pausing and mRNA instability.
A good optimization tool will screen for these motifs, but it is prudent to manually verify the output. The Start Codon and Stop Codon context should also be checked: the start codon should be preceded by a strong Shine-Dalgarno sequence (AGGAGG) at an optimal spacing of 5–9 nucleotides, and the stop codon should be followed by a transcription terminator.
Summary and Practical Recommendations
Step-by-Step Workflow
- Assess the native gene: Calculate the CAI of the original sequence using the E. coli codon usage table. If the CAI is above 0.7 and the gene is not AT-rich, consider expressing it without optimization.
- Choose an optimization strategy: For most proteins, a balanced optimization targeting a CAI of 0.85–0.95 is appropriate. For proteins known to be prone to misfolding, consider harmonized optimization that preserves the relative codon usage pattern of the native gene.
- Select a tool: Use a tool that allows constraints on GC content, mRNA structure, and restriction sites. Compare outputs from two different tools to identify consensus regions.
- Check the output manually: Verify the absence of SD-like sequences, restriction sites, and extreme GC content. Predict the mRNA secondary structure at the 5' end.
- Synthesize and clone: Order the optimized gene with flanking restriction sites or Gibson homology arms. Clone into an E. coli Expression System such as pET-28a.
- Test expression: Perform a small-scale expression test with a time course (0, 1, 2, 4 hours post-induction). Analyze both soluble and insoluble fractions.
- Iterate if needed: If expression is poor, test alternative optimization strategies, different expression strains, or lower induction temperatures.
Integration with Other Expression Strategies
Codon optimization should be integrated with other expression-enhancing strategies. Fusion partners such as maltose-binding protein (MBP) or glutathione S-transferase (GST) can improve solubility. Co-expression of chaperones (GroEL/GroES, DnaK/DnaJ/GrpE) can assist folding. The choice of E. coli strain matters: BL21(DE3) is standard, but Rosetta strains carry extra tRNA genes for rare codons, and C43(DE3) is better for membrane proteins.
For proteins that remain insoluble after optimization, consider refolding from inclusion bodies. This requires denaturation in 6 M guanidine-HCl or 8 M urea, followed by refolding by dialysis or dilution. The refolding buffer typically contains 50 mM Tris-HCl, pH 8.0, 100 mM NaCl, and a redox pair such as 1 mM reduced glutathione and 0.1 mM oxidized glutathione.
The E. coli Protein Expression resource provides a comprehensive overview of these complementary strategies. For a deeper understanding of the underlying translation mechanisms, the Codon Anticodon interaction is the fundamental determinant of codon optimality.
Frequently Asked Questions
What is codon optimization for E. coli?
Codon optimization is the redesign of a coding sequence to replace codons that are rare in E. coli with synonymous codons that are more frequently used, while preserving the amino acid sequence. The goal is to improve translation efficiency, mRNA stability, and ultimately protein yield. The process also typically addresses secondary factors such as GC content, mRNA secondary structure, and the removal of unintended regulatory motifs.
Does codon optimization always increase protein yield in E. coli?
No. Codon optimization increases yield when the native gene contains rare codons that slow translation or destabilize mRNA. However, if the native gene is already well-adapted, or if the protein's folding depends on specific translation kinetics, optimization can be neutral or even detrimental. Over-optimization can reduce yield by creating mRNA secondary structures or by disrupting co-translational folding.
What is the best codon optimization tool for E. coli?
There is no single best tool. IDT and SnapGene are user-friendly and suitable for most applications. JCAT offers more sophisticated multi-objective optimization. For proteins prone to misfolding, tools that preserve codon pair bias, such as Optimizer, may be preferable. The choice should be guided by the specific constraints of your project, and it is advisable to compare outputs from at least two tools.
How does codon usage affect translation in E. coli?
Codon usage affects translation at multiple levels. During elongation, codons with abundant cognate tRNAs are decoded rapidly, while rare codons cause ribosome pausing. Pausing can trigger tmRNA-mediated degradation of the nascent protein. Codon usage also affects mRNA stability, because ribosome density protects mRNA from RNase E cleavage. Finally, translation kinetics influence co-translational folding, with pauses at domain boundaries facilitating proper folding.
What is codon adaptation index (CAI) and how is it used?
The Codon Adaptation Index is a metric that quantifies how well a coding sequence matches the codon usage of a reference set of highly expressed genes. It is calculated as the geometric mean of relative synonymous codon usage values, and ranges from 0 to 1. A CAI above 0.8 is generally considered good for E. coli expression. CAI is used to assess whether a gene needs optimization and to evaluate the output of optimization tools.
Should I avoid rare codons in E. coli?
Rare codons should be avoided in clusters, but isolated rare codons are generally tolerable. Clusters of three or more rare codons, particularly AGG/AGA (Arg), ATA (Ile), and CTA (Leu), can cause ribosome stalling and reduced yield. However, some proteins require rare codons at specific positions for proper folding. If the native gene expresses acceptably, there is no need to eliminate all rare codons.
Can codon optimization affect mRNA stability?
Yes. Codon optimization affects mRNA stability through its influence on ribosome density. Transcripts with optimal codons are translated rapidly and maintain high ribosome density, which protects them from RNase E cleavage. Transcripts with many non-optimal codons have gaps in ribosome coverage and are degraded more quickly. Optimization therefore typically increases mRNA half-life, which contributes to higher protein yield.
What are common mistakes in codon optimization?
The most common mistakes are over-optimization to a CAI above 0.95, ignoring mRNA secondary structure at the 5' end, and failing to check for unintended motifs such as internal Shine-Dalgarno sequences, restriction sites, and Rho-independent terminators. Another frequent error is assuming that codon optimization alone will solve all expression problems, when factors such as protein toxicity, codon pair bias, and co-translational folding are equally important.
Key Takeaways
- Codon optimization redesigns a coding sequence to match E. coli tRNA availability and codon usage bias, improving translation efficiency and mRNA stability.
- The Codon Adaptation Index (CAI) is the standard metric; a target of 0.85–0.95 is generally optimal, and chasing a CAI of 1.0 often causes problems.
- Translation kinetics matter: rare codon clusters cause ribosome stalling and tmRNA-mediated degradation, but programmed pauses at domain boundaries can be important for folding.
- mRNA secondary structure at the 5' end is a critical constraint; the ribosome binding site and start codon must remain accessible.
- Codon pair bias and GC content should be balanced, not maximized, to avoid unintended regulatory motifs and extreme sequence composition.
- Optimization is not always necessary or beneficial; assess the native gene's CAI first, and consider harmonized optimization for proteins prone to misfolding.
- Always validate optimized genes experimentally with small-scale expression tests, and combine optimization with appropriate expression strains, fusion partners, and induction conditions.
Further Reading
- Paremskaia AI et al. Codon-optimization in gene therapy: promises, prospects and challenges. Frontiers in bioengineering and biotechnology. 2024. PubMed 38605988
- Jenkins MC et al. Effects of codon optimization on expression in Escherichia coli of protein-coding DNA sequences from the protozoan Eimeria. Journal of microbiological methods. 2023. PubMed 37271377
- Gao CY et al. Codon optimization enhances the expression of porcine β-defensin-2 in Escherichia coli. Genetics and molecular research : GMR. 2015. PubMed 25966273
- Nieuwkoop T, Claassens NJ, van der Oost J. Improved protein production and codon optimization analyses in Escherichia coli by bicistronic design. Microbial biotechnology. 2019. PubMed 30484964
- Wong DPH et al. OPT: Codon optimize gene sequences for E. coli protein overexpression. Journal of molecular biology. 2025. PubMed 40133777
- Song H et al. Codon optimization enhances protein expression of Bombyx mori nucleopolyhedrovirus DNA polymerase in E. coli. Current microbiology. 2014. PubMed 24129839