Custom Recombinant Protein Expression: A Practical Guide
By Dr. Zubair Khalid, DVM, MS, PhD ·

Introduction to Custom Recombinant Protein Expression
Recombinant protein expression is the process by which a protein encoded by a cloned gene is produced in a heterologous host—that is, a host organism different from the one in which the gene naturally occurs. When the gene of interest is supplied by the researcher rather than selected from a pre-existing library, the process is referred to as custom recombinant protein expression. This approach allows you to produce virtually any protein, from a human enzyme to a bacterial toxin, in a controlled laboratory setting, and to obtain it in quantities far exceeding what could be purified from the native source.
The purpose of custom expression is threefold: to obtain sufficient protein for structural studies (such as X-ray crystallography or cryo-electron microscopy), to generate material for biochemical characterization and activity assays, and to produce antigens or therapeutic proteins for applied use. The entire workflow, from gene to purified protein, follows a logical sequence of decisions and manipulations.
The overall workflow begins with the acquisition of the gene of interest, typically by PCR amplification from genomic DNA or cDNA, or by chemical synthesis. This gene is then inserted into an expression vector—a plasmid engineered to drive transcription and translation in the chosen host. The resulting construct is introduced into the host cells, which are cultured under conditions that induce protein production. Finally, the protein is extracted from the cells and purified to the required degree. Each of these steps presents choices that materially affect the outcome, and this guide walks through those choices in the order you will encounter them.
What is Recombinant Protein Expression?
Recombinant protein expression exploits the universal nature of the genetic code. A gene from any organism, once placed under the control of regulatory elements recognized by the host's transcriptional and translational machinery, will direct the synthesis of the same polypeptide sequence in the host as it does in the native organism. The host then performs the essential functions of protein synthesis: transcription of DNA to mRNA, translation of mRNA to polypeptide, and—depending on the host—post-translational modifications such as glycosylation, phosphorylation, or disulfide bond formation.
The key distinction between recombinant expression and native purification is control. In recombinant expression, you decide which gene is expressed, in which organism, at what time, and to what level. You can attach purification tags, introduce mutations, or express only a domain of a larger protein. This control is what makes the approach so powerful for both basic research and biotechnology.
Why Custom Expression?
You might ask why one would not simply purify the protein from its natural source. The answer is often practical: the native source may be scarce, the protein may be present at very low abundance, or the source organism may be difficult to culture. A human protein, for example, might be present in nanogram quantities per gram of tissue, whereas a recombinant bacterial system can produce milligrams per liter of culture. Custom expression also enables the production of proteins that are toxic to the native organism, or proteins with engineered modifications such as isotopic labeling for NMR studies.
The choice of expression system is the first major decision, and it is dictated by the properties of the protein you wish to produce. The following section compares the most commonly used hosts.
Choosing the Right Expression Host
The host organism is the factory in which your protein is made. Its choice determines the yield, the solubility, the post-translational modifications, and the cost of production. There is no single "best" host; there is only the best host for a given protein.
Bacterial Systems
Escherichia coli is the workhorse of recombinant protein expression, and for good reason. It grows rapidly to high cell density in inexpensive media, its genetics are thoroughly characterized, and a vast array of plasmids and strains are available. For proteins that do not require glycosylation and that fold correctly in the reducing environment of the bacterial cytoplasm, E. coli can produce gram-per-liter yields.
The most widely used E. coli expression system is based on the T7 RNA polymerase, driven by the pET plasmid series. In this system, the gene of interest is cloned downstream of a T7 promoter. The host strain (such as BL21(DE3)) carries a chromosomal copy of the T7 RNA polymerase gene under the control of the lacUV5 promoter. Addition of isopropyl β-D-1-thiogalactopyranoside (IPTG) induces the lacUV5 promoter, producing T7 RNA polymerase, which then transcribes the gene of interest at high levels.
However, E. coli has limitations. It cannot perform N-linked glycosylation, it lacks the machinery for complex disulfide bond formation in the cytoplasm, and many eukaryotic proteins misfold and aggregate when overexpressed in bacteria. For these proteins, eukaryotic systems are required.
Other bacterial hosts include Bacillus subtilis, which has a robust secretory pathway and is used for industrial enzyme production, and Lactococcus lactis, which is food-grade and useful for producing proteins intended for human consumption.
Eukaryotic Systems
When post-translational modifications are essential for protein function, you must turn to eukaryotic hosts.
Yeast (Saccharomyces cerevisiae and Pichia pastoris) offers a middle ground. Yeast grows quickly in inexpensive media, and P. pastoris can achieve very high cell densities. Yeast performs core eukaryotic modifications, including signal peptide processing, disulfide bond formation, and some glycosylation. However, yeast glycosylation patterns differ from those of mammals, adding high-mannose structures that can be immunogenic in humans. For proteins that tolerate these modifications, yeast is an excellent, cost-effective choice.
Insect cells, typically derived from Spodoptera frugiperda (Sf9 or Sf21) or Trichoplusia ni (High Five), are infected with a recombinant baculovirus carrying the gene of interest. These cells perform more complex post-translational modifications than yeast, including glycosylation that is closer to mammalian patterns. Baculovirus expression is particularly useful for large, multi-domain proteins and for proteins that are toxic to bacteria. The downside is that insect cell culture requires more specialized equipment and media, and the baculovirus infection eventually lyses the cells, limiting the production window.
Mammalian cells, such as Chinese hamster ovary (CHO) cells or human embryonic kidney (HEK) 293 cells, provide the most authentic post-translational modifications, including human-type glycosylation. These systems are the standard for therapeutic protein production, such as antibodies. However, mammalian cell culture is slow, expensive, and technically demanding. For most academic research applications, mammalian expression is reserved for proteins that cannot be produced correctly in any other system.
Factors Influencing Host Choice
The decision matrix below summarizes the key considerations.
| Host | Typical Yield | Glycosylation | Disulfide Bonds | Cost | Time | Best For |
|---|---|---|---|---|---|---|
| E. coli | mg to g/L | None | Poor in cytoplasm | Low | Days | Cytosolic proteins, unmodified proteins, high-throughput screens |
| Yeast (P. pastoris) | mg to g/L | High-mannose | Yes | Low | Days to weeks | Secreted proteins, proteins tolerating non-mammalian glycosylation |
| Insect cells | mg/L | Simple, near-mammalian | Yes | Medium | Weeks | Multi-domain proteins, membrane proteins |
| Mammalian cells | µg to mg/L | Human-type | Yes | High | Weeks to months | Therapeutic proteins, antibodies, proteins requiring complex modifications |
For a deeper comparison of available platforms, see the Recombinant Protein Expression System resource. If you are considering outsourcing the work, Contract Recombinant Protein Expression services can handle host selection, construct design, and purification on your behalf.
Designing the Expression Construct
Once the host is chosen, the next task is to design the expression construct—the plasmid that will carry your gene into the host and direct its expression. The construct must contain several essential elements, each of which can be optimized.
Promoters and Induction
The promoter is the DNA sequence that recruits RNA polymerase to initiate transcription. For inducible expression, the promoter is kept off during cell growth and turned on when you are ready to produce protein. This is critical because high-level expression of a foreign protein often imposes a metabolic burden on the host, slowing growth and selecting for plasmid loss or mutations.
In E. coli, the most common inducible promoters are:
- T7 promoter (in pET vectors): Induced by IPTG, which activates T7 RNA polymerase. Very strong, producing up to 50% of total cellular protein.
- araBAD promoter (pBAD vectors): Induced by L-arabinose. Tightly regulated and titratable—expression level is proportional to arabinose concentration.
- rhaBAD promoter: Induced by L-rhamnose. Similar to araBAD but with even tighter basal repression.
- trc and tac promoters: Hybrid promoters induced by IPTG. Moderate strength, useful for proteins that are toxic when overexpressed.
In yeast, the most common inducible promoter is the AOX1 promoter in P. pastoris, which is induced by methanol. In mammalian cells, the CMV promoter is constitutively active, while the tetracycline-inducible system (Tet-On/Tet-Off) allows regulated expression.
The choice of promoter affects not only the yield but also the kinetics of induction. A strong promoter like T7 produces protein rapidly, but this can overwhelm the folding machinery and lead to inclusion bodies. A weaker promoter may produce less protein per cell, but a larger fraction of it may be soluble and correctly folded.
Fusion Tags and Their Functions
A fusion tag is a peptide or protein sequence attached to your protein of interest. Tags serve two primary purposes: purification and solubility enhancement.
Affinity tags enable one-step purification. The most common is the polyhistidine tag (His-tag), typically six consecutive histidine residues. The His-tag binds to immobilized nickel or cobalt ions with micromolar affinity, allowing the tagged protein to be captured on a metal-affinity resin while contaminating proteins are washed away. Elution is achieved with imidazole, which competes for the metal-binding sites. The His-tag is small (approximately 0.8 kDa), rarely affects protein structure, and can be placed at either the N- or C-terminus.
The glutathione S-transferase (GST) tag (26 kDa) is a larger tag that binds to glutathione-agarose resin. GST tags often improve solubility of the fused protein, but the larger size increases the chance of interfering with folding. The maltose-binding protein (MBP) tag (42 kDa) is even more effective at enhancing solubility, but it is large enough that it must almost always be removed after purification.
Solubility tags are used when the protein of interest tends to aggregate. In addition to GST and MBP, the small ubiquitin-like modifier (SUMO) tag is popular because it both enhances solubility and can be cleaved with high specificity by the SUMO protease Ulp1, leaving no residual amino acids on the protein of interest.
Detection tags, such as the FLAG tag (DYKDDDDK) or the c-Myc tag (EQKLISEEDL), are short peptides recognized by commercial antibodies. They are useful for Western blotting and immunoprecipitation but are not typically used for purification.
Tags are usually separated from the protein of interest by a protease cleavage site, such as the recognition sequence for tobacco etch virus (TEV) protease (ENLYFQG) or PreScission protease (LEVLFQGP). Cleavage occurs at a defined position, allowing the tag to be removed after purification.
Codon Optimization
The genetic code is degenerate: most amino acids are encoded by multiple codons. Different organisms have different preferences for which codon is used for a given amino acid, and these preferences correlate with the abundance of the corresponding tRNA. If your gene contains codons that are rare in the host, the ribosome may stall at those positions, leading to premature termination, reduced yield, or misfolding.
Codon optimization is the process of altering the DNA sequence of the gene to match the codon usage bias of the host, without changing the amino acid sequence. This is typically done by commercial gene synthesis services, which can also remove problematic sequences such as internal restriction sites, repetitive sequences, and mRNA secondary structures.
For example, human genes expressed in E. coli often contain codons for arginine (AGA, AGG) and proline (CCC) that are rare in bacteria. The E. coli strain Rosetta (BL21(DE3)pLysS) carries extra copies of the tRNA genes for these rare codons, providing a workaround if codon optimization is not possible. However, for optimal expression, especially of large proteins, codon optimization is strongly recommended.
Cloning Strategies for Custom Expression
With the vector design in hand, the next step is to physically insert the gene of interest into the plasmid. Several methods are available, each with advantages and limitations.
Traditional Restriction Cloning
Restriction cloning is the classical approach. The gene of interest is amplified by PCR with primers that incorporate restriction enzyme recognition sites at the 5' and 3' ends. The PCR product and the vector are both digested with the same restriction enzymes, producing compatible cohesive ends. The digested fragments are then ligated together with T4 DNA ligase.
The key requirement is that the restriction sites must be absent from the interior of the gene of interest. This can be checked by sequence analysis before designing the primers. Common choices include NdeI (CATATG), which conveniently includes the ATG start codon, and XhoI (CTCGAG), BamHI (GGATCC), and EcoRI (GAATTC).
The main limitation of restriction cloning is that it leaves extra nucleotides at the junction between the vector and the insert—the restriction sites themselves. If these nucleotides encode amino acids, they will appear in the final protein, potentially affecting its structure or function. For this reason, many researchers prefer seamless cloning methods.
Seamless Cloning Techniques
Gibson assembly is a powerful method that joins multiple DNA fragments in a single isothermal reaction. The fragments are designed with overlapping ends of 20-40 base pairs. The reaction contains three enzymes: a 5'→3' exonuclease (T5 exonuclease) that chews back the 5' ends to create single-stranded overhangs, a DNA polymerase (Phusion) that fills in the gaps, and a DNA ligase (Taq ligase) that seals the nicks. The result is a seamless junction with no extra nucleotides.
Gibson assembly is ideal for cloning large genes, for combining multiple fragments (such as a promoter, gene, and terminator), and for site-directed mutagenesis. The reaction is performed at 50°C for 1 hour and requires only that the overlapping sequences be designed correctly.
Ligation-independent cloning (LIC) uses the 3'→5' exonuclease activity of T4 DNA polymerase to create long, complementary single-stranded overhangs. The vector is linearized and treated with T4 DNA polymerase in the presence of only one dNTP, causing the polymerase to chew back until it encounters a nucleotide matching the provided dNTP. The insert is treated similarly with a different dNTP. The two fragments are then annealed, and the nicks are repaired upon transformation into E. coli. LIC requires no ligase and no restriction enzymes, and it produces seamless junctions.
Golden Gate assembly uses type IIS restriction enzymes, such as BsaI or BsmBI, which cut outside their recognition sequence. This allows the recognition site to be removed during digestion, leaving custom overhangs that direct the assembly of multiple fragments in a defined order. Golden Gate is particularly useful for modular cloning, where standardized parts are assembled into larger constructs.
Transformation and Expression Screening
Once the construct is assembled, it must be introduced into the host cells. This process is called transformation for bacteria and yeast, and transfection for mammalian cells.
Transformation Methods
For E. coli, the two standard methods are chemical transformation and electroporation.
Chemical transformation uses calcium chloride to make the cells competent—that is, permeable to DNA. Cells are incubated on ice with the plasmid, heat-shocked at 42°C for 45 seconds, and then allowed to recover in rich medium before plating on selective agar. This method is simple and inexpensive, with efficiencies of 10⁶–10⁸ colony-forming units per microgram of DNA.
Electroporation uses a brief electrical pulse to create transient pores in the cell membrane. Electrocompetent cells are mixed with the plasmid and subjected to a high-voltage pulse (typically 1.8 kV for 0.1 cm cuvettes). Electroporation is more efficient than chemical transformation (10⁹–10¹⁰ CFU/µg) and is preferred for large plasmids or when high efficiency is critical.
After transformation, cells are plated on agar containing the appropriate antibiotic—ampicillin (100 µg/mL), kanamycin (50 µg/mL), or chloramphenicol (34 µg/mL)—to select for cells that have acquired the plasmid. Individual colonies are then picked and grown in liquid culture for screening.
Screening for Expression
The first screen is colony PCR, which confirms that the insert is present in the plasmid. A small amount of a colony is added directly to a PCR reaction with primers that flank the insertion site. The size of the PCR product indicates whether the insert is present and in the correct orientation.
The second screen is small-scale expression testing. Several colonies are grown in 5 mL cultures, induced with IPTG (or the appropriate inducer), and harvested after 3-4 hours. The cells are lysed, and the total protein is analyzed by SDS-PAGE. A band of the expected molecular weight that appears after induction, and is absent in the uninduced control, indicates successful expression. This step also reveals whether the protein is soluble or present in inclusion bodies, which is visible as a difference between the soluble and insoluble fractions.
For high-throughput screening of many constructs or conditions, see the Recombinant Protein Lab protocols, which describe automated systems for parallel expression testing.
Optimizing Expression Conditions
Even when a construct expresses protein, the yield and solubility may be suboptimal. Optimization involves systematically varying the conditions of induction and growth.
Induction Optimization
The key variables are inducer concentration, temperature, and time.
For IPTG-inducible systems, the standard concentration is 0.5–1 mM. However, high IPTG concentrations force rapid, high-level expression that can overwhelm the folding machinery. Reducing IPTG to 0.1–0.2 mM often improves solubility with only a modest reduction in total yield. This is because slower production gives the chaperones time to fold the protein correctly.
Temperature is the most powerful variable. At 37°C, protein synthesis is fast, but so is aggregation. Lowering the temperature to 25°C or even 16°C slows translation, which often dramatically improves solubility. The trade-off is that lower temperatures require longer induction times—overnight induction at 16°C is common.
Time is the third variable. For many proteins, the yield plateaus after 3-4 hours at 37°C, after which the protein may begin to degrade. For slow, low-temperature inductions, 16-20 hours is typical. The optimal time must be determined empirically for each protein.
Solubility Enhancement Strategies
If the protein is expressed but insoluble, several strategies can be tried:
- Lower the temperature to 16-25°C during induction.
- Reduce IPTG concentration to 0.05–0.2 mM.
- Use a different strain that provides chaperones or rare tRNAs. Strains like BL21(DE3)pLysS reduce basal expression, while Rosetta(DE3) supplies rare tRNAs.
- Fuse a solubility tag such as MBP or SUMO to the N-terminus.
- Co-express molecular chaperones such as GroEL/GroES or DnaK/DnaJ, which assist in protein folding.
- Change the growth medium to a rich medium like Terrific Broth, which supports higher cell density and may improve folding.
- Modify the induction strategy—for example, use auto-induction media, which gradually induce expression as the cells deplete glucose, rather than a sudden IPTG shock.
For a systematic approach to solubility problems, the Recombinant Protein Solubility Expression guide provides a decision tree for troubleshooting.
Protein Purification Strategies
After expression, the protein must be extracted from the cells and purified. The purification strategy depends on the tag used and the required purity.
Affinity Chromatography
Affinity chromatography is the first step in virtually all recombinant protein purifications. It exploits the specific, reversible interaction between the tag and a ligand immobilized on a resin.
His-tag purification uses resins charged with Ni²⁺ or Co²⁺ ions, such as Ni-NTA (nitrilotriacetic acid) agarose. The workflow is:
- Lyse the cells by sonication or French press in a lysis buffer containing 20–50 mM sodium phosphate (pH 7.4–8.0), 300–500 mM NaCl, and 10–20 mM imidazole. The NaCl prevents non-specific ionic interactions, and the low imidazole concentration reduces non-specific binding of host proteins.
- Clarify the lysate by centrifugation at 20,000 × g for 30 minutes at 4°C.
- Incubate the supernatant with Ni-NTA resin for 30-60 minutes at 4°C with gentle agitation.
- Wash the resin with buffer containing 20–50 mM imidazole to remove weakly bound contaminants.
- Elute the protein with buffer containing 250–500 mM imidazole, which competes with the His-tag for the nickel binding sites.
The eluted protein is typically 80-90% pure after a single affinity step.
GST purification uses glutathione-agarose resin. The GST-tagged protein binds to the resin, is washed, and is eluted with 10–20 mM reduced glutathione in 50 mM Tris-HCl (pH 8.0). GST purification often yields cleaner protein than His-tag purification, but the resin is more expensive.
MBP purification uses amylose resin, and elution is achieved with 10 mM maltose. MBP-tagged proteins often require gentler elution conditions and are less prone to non-specific binding.
Tag Removal and Further Purification
If the tag must be removed, the eluted protein is incubated with the appropriate protease. TEV protease is commonly used because it is highly specific and active at 4°C. The cleavage reaction is typically performed overnight at 4°C in a buffer containing 1 mM DTT (TEV requires a reducing agent) and 0.5 mM EDTA.
After cleavage, the tag and the protease must be separated from the protein of interest. This is often done by a second pass over the affinity resin: the His-tagged TEV protease and the cleaved His-tag bind to the Ni-NTA resin, while the untagged protein flows through. Alternatively, size-exclusion chromatography (gel filtration) can separate the protein from the tag based on size.
Size-exclusion chromatography (SEC) is the standard polishing step. The protein is loaded onto a column packed with porous beads; smaller proteins enter the pores and are retarded, while larger proteins pass through more quickly. SEC also exchanges the buffer, which is useful for removing imidazole or glutathione.
Ion-exchange chromatography separates proteins based on surface charge. A cation-exchange column (e.g., SP-Sepharose) binds positively charged proteins, while an anion-exchange column (e.g., Q-Sepharose) binds negatively charged proteins. Elution is achieved with a salt gradient, typically 0–1 M NaCl.
Analyzing and Validating the Recombinant Protein
Purification is not the end of the process. The protein must be confirmed to be the correct product, pure, and functional.
SDS-PAGE and Western Blotting
SDS-PAGE (sodium dodecyl sulfate-polyacrylamide gel electrophoresis) is the standard method for assessing purity and molecular weight. The protein is denatured with SDS, which coats the polypeptide with a uniform negative charge, and separated by size through a polyacrylamide gel. After electrophoresis, the gel is stained with Coomassie Blue, which detects protein bands at a sensitivity of approximately 100 ng.
A single band at the expected molecular weight indicates high purity. However, the presence of additional bands may indicate degradation products, co-purifying contaminants, or incomplete removal of the tag.
Western blotting confirms the identity of the protein. The proteins from an SDS-PAGE gel are transferred to a nitrocellulose or PVDF membrane, which is then probed with an antibody specific to the protein or to its tag. Detection is achieved with a secondary antibody conjugated to horseradish peroxidase (HRP), which catalyzes a chemiluminescent reaction that exposes X-ray film or is detected by a digital imager. Western blotting can detect as little as 1 ng of protein and is essential when the protein is not visible by Coomassie staining.
Functional Assays
The ultimate validation is a functional assay. The nature of this assay depends on the protein. For an enzyme, you would measure its catalytic activity—for example, the conversion of a substrate to a product, monitored by spectrophotometry or chromatography. For a binding protein, you would measure its affinity for a ligand using techniques such as surface plasmon resonance or isothermal titration calorimetry. For a structural protein, you might assess its ability to polymerize or to interact with other proteins.
Mass spectrometry provides definitive confirmation of identity. The purified protein is digested with trypsin, and the resulting peptides are analyzed by LC-MS/MS. The peptide masses are compared against the predicted sequence, confirming that the protein has the correct primary structure and that no unintended mutations are present.
Common Pitfalls and Troubleshooting
Even with careful planning, recombinant protein expression frequently fails. The most common problems, and their solutions, are described below.
Inclusion Bodies and Solubility
Inclusion bodies are dense, insoluble aggregates of misfolded protein that form when the rate of synthesis exceeds the capacity of the folding machinery. They are visible as refractile bodies under the microscope and pellet at low speed during centrifugation.
If the protein is in inclusion bodies, you have two options: refold the protein, or prevent the aggregation in the first place.
Refolding involves solubilizing the inclusion bodies in a strong denaturant (6–8 M urea or 6 M guanidine hydrochloride), then slowly removing the denaturant by dialysis or dilution to allow the protein to fold. Refolding is often inefficient and can produce protein with incorrect disulfide bonds or non-native conformations, but it is sometimes the only option for proteins that are toxic to the host.
Prevention is preferable. Lower the temperature, reduce the inducer concentration, use a weaker promoter, or fuse a solubility tag. The Recombinant Protein Solubility Expression resource provides a detailed troubleshooting flowchart.
Proteolysis and Stability
Many recombinant proteins are degraded by host proteases. This is evident as multiple lower-molecular-weight bands on SDS-PAGE, or as a gradual loss of full-length protein over time.
Solutions include:
- Use protease-deficient strains. E. coli strains such as BL21(DE3) are deficient in Lon and OmpT proteases. For particularly sensitive proteins, strains like Rosetta-gami or Lemo21(DE3) offer additional advantages.
- Add protease inhibitors to the lysis buffer. A cocktail containing PMSF (1 mM), leupeptin (1 µg/mL), and pepstatin (1 µg/mL) is standard.
- Work at 4°C during all purification steps.
- Purify quickly. Minimize the time between cell lysis and affinity capture.
- Fuse a large tag such as MBP, which can protect the protein of interest from proteolysis.
Low Expression Levels
If the protein is not expressed at all, or expressed at very low levels, the cause is often one of the following:
- The gene is not in the correct reading frame relative to the start codon. This is confirmed by sequencing the construct.
- The promoter is not recognized by the host. For example, a mammalian promoter will not work in E. coli.
- The mRNA is unstable or poorly translated. This can be addressed by codon optimization, by removing secondary structure at the 5' end, or by changing the sequence around the ribosome binding site.
- The protein is toxic to the host. Cells carrying the plasmid may grow slowly or die upon induction. In this case, use a tightly regulated promoter (such as araBAD or rhaBAD), induce at low cell density, or use a strain that suppresses basal expression.
- The plasmid is lost. If the antibiotic is degraded or the selective pressure is insufficient, cells may lose the plasmid. Ensure the antibiotic is fresh and used at the correct concentration.
For a comprehensive overview of troubleshooting strategies, the Recombinant Technology for Protein Expression article covers the full range of technical approaches.
Frequently Asked Questions
What is custom recombinant protein expression?
Custom recombinant protein expression is the production of a specific protein in a heterologous host, where the gene encoding the protein is provided by the researcher. The gene is cloned into an expression vector, introduced into a host organism, and the host's cellular machinery is used to synthesize the protein. The term "custom" distinguishes this from using pre-existing expression libraries or off-the-shelf proteins.
Which host is best for recombinant protein expression?
There is no universal answer. E. coli is the best starting point for most proteins because it is fast, inexpensive, and high-yielding. However, if the protein requires glycosylation or complex disulfide bond formation, a eukaryotic host such as yeast, insect cells, or mammalian cells is necessary. The choice depends on the protein's size, complexity, and required post-translational modifications.
What are fusion tags and why are they used?
Fusion tags are additional peptide or protein sequences attached to the protein of interest. They are used for three main purposes: purification (e.g., His-tag, GST-tag), solubility enhancement (e.g., MBP, SUMO), and detection (e.g., FLAG, c-Myc). Tags are typically removed after purification using a site-specific protease.
How do I choose the right expression vector?
The vector must be compatible with the host, contain an appropriate promoter, and carry a selectable marker (antibiotic resistance gene). For E. coli, pET vectors (T7 promoter) are standard. For yeast, pPICZ vectors (AOX1 promoter) are common. The vector should also have a multiple cloning site with restriction sites that are absent from your gene, and an appropriate fusion tag for purification.
Why is codon optimization important?
Codon optimization adjusts the DNA sequence of the gene to match the codon usage preferences of the host. If the gene contains codons that are rare in the host, translation may stall, leading to low yield, truncated products, or misfolding. Optimized genes are usually produced by commercial gene synthesis.
What causes inclusion bodies and how can I avoid them?
Inclusion bodies form when the protein is produced faster than it can fold, causing partially folded intermediates to aggregate. They can be avoided by lowering the induction temperature, reducing the inducer concentration, using a weaker promoter, or fusing a solubility tag. If inclusion bodies form, the protein can sometimes be refolded in vitro.
How do I purify a His-tagged protein?
A His-tagged protein is purified by immobilized metal affinity chromatography (IMAC). The cell lysate is incubated with Ni-NTA resin, which binds the His-tag. The resin is washed with buffer containing low imidazole (20–50 mM), and the protein is eluted with high imidazole (250–500 mM). The purified protein can then be dialyzed to remove the imidazole.
What are common pitfalls in recombinant protein expression?
The most common pitfalls are low or absent expression, protein insolubility (inclusion bodies), proteolytic degradation, and contamination during purification. Each has specific troubleshooting strategies, as described in the Common Pitfalls section above.
Key Takeaways
- Custom recombinant protein expression is a multi-step workflow: gene acquisition, vector construction, host transformation, expression, purification, and validation.
- The choice of expression host is the most consequential decision; E. coli is the default, but eukaryotic hosts are required for proteins needing post-translational modifications.
- Expression constructs require a promoter, ribosome binding site, selectable marker, and often a fusion tag for purification and solubility.
- Seamless cloning methods (Gibson assembly, LIC, Golden Gate) are preferred over traditional restriction cloning because they leave no extra amino acids.
- Expression conditions—temperature, inducer concentration, and time—are the primary variables for optimizing yield and solubility.
- Affinity purification using His-tags is the standard first step; tag removal and size-exclusion chromatography provide final purity.
- Validation requires SDS-PAGE, Western blotting, and a functional assay to confirm identity, purity, and activity.
- Inclusion bodies, proteolysis, and low expression are the most common failures, and each has established troubleshooting strategies. For complex projects, Contract Recombinant Protein Expression services and Express Recombinant Protein platforms can provide specialized expertise.
Further Reading
- Hayat SMG et al. Recombinant Protein Expression in Escherichia coli (E.coli): What We Need to Know. Current pharmaceutical design. 2018. PubMed 29384059
- Ritacco FV, Wu Y, Khetan A. Cell culture media for recombinant protein expression in Chinese hamster ovary (CHO) cells: History, key components, and optimization strategies. Biotechnology progress. 2018. PubMed 30290072
- Bazaz M et al. Recent developments in miRNA based recombinant protein expression in CHO. Biotechnology letters. 2022. PubMed 35507207
- Andersen DC, Krummen L. Recombinant protein expression for therapeutic applications. Current opinion in biotechnology. 2002. PubMed 1195056100300-2)
- Marschall L, Sagmeister P, Herwig C. Tunable recombinant protein expression in E. coli: enabler for continuous processing?. Applied microbiology and biotechnology. 2016. PubMed 27170324
- Zhang HJ et al. Effect of Apoptosis and Autophagy on Recombinant Protein Expression in Chinese Hamster Ovary Cells. Biotechnology journal. 2025. PubMed 40619711