Protein Expression and Purification: A Practical Guide
By Dr. Zubair Khalid, DVM, MS, PhD ·

Introduction to Protein Expression and Purification
The central goal of protein expression and purification is to obtain a single protein of interest in sufficient quantity, at high purity, and in a conformation that retains its biological activity. This undertaking is foundational to nearly every branch of molecular biology, biochemistry, and structural biology. Whether you are preparing an enzyme for kinetic assays, a receptor for ligand-binding studies, or an antigen for antibody production, the quality of your downstream data is directly limited by the quality of your protein preparation.
Why purify proteins?
Purification is necessary because no expression system produces a protein in isolation. Even in the most favorable cases, the target protein constitutes only a fraction of the total cellular protein. In Escherichia coli, a highly expressed recombinant protein might represent 10–40% of total soluble protein, but that still leaves the majority of the lysate as contaminating host proteins, nucleic acids, lipids, and other cellular components. Many applications—X-ray crystallography, cryo-electron microscopy, nuclear magnetic resonance spectroscopy, and quantitative enzyme assays—require protein that is >95% pure and free of interfering contaminants such as proteases, nucleases, or endotoxins.
Beyond purity, the protein must be functional. A protein that is pure but misfolded, aggregated, or improperly modified is useless for most applications. This dual requirement—purity and functionality—drives every decision in the workflow, from the choice of expression host to the final polishing step.
Overview of the expression-purification pipeline
The complete workflow can be divided into six stages:
- Gene acquisition and construct design — Codon optimization, cloning into an expression vector with appropriate regulatory elements and affinity tags.
- Expression — Transformation into the host, growth to appropriate cell density, and induction of protein production.
- Harvest and lysis — Cell collection, disruption, and preparation of a clarified lysate.
- Capture — A high-capacity, high-selectivity step (usually affinity chromatography) that isolates the target from the bulk of contaminants.
- Polishing — Additional chromatography steps (ion exchange, size exclusion) that remove residual contaminants and aggregate species.
- Characterization — Assessment of purity, identity, and activity using biochemical and biophysical methods.
Each stage presents distinct challenges, and the choices made at one stage propagate through all subsequent stages. A poorly designed construct, for example, cannot be rescued by even the most sophisticated purification strategy. This guide walks through each stage with mechanistic detail and practical recommendations.
Choosing an Expression System
The expression host determines the chemical environment in which your protein is synthesized, folded, and post-translationally modified. There is no universal best system; the optimal choice depends on the protein's origin, size, complexity, and intended use.
E. coli: advantages and limitations
E. coli remains the workhorse of recombinant protein production for good reason. It grows rapidly (doubling time ~20 minutes in rich media), reaches high cell densities, and is inexpensive to culture. The genetics are well understood, and a vast toolkit of plasmids, strains, and induction systems exists. For proteins that do not require glycosylation or other eukaryotic modifications, E. coli can produce gram-per-liter quantities in fermenters and tens of milligrams per liter in shake flasks.
The limitations are significant, however. E. coli lacks the machinery for N-linked glycosylation, disulfide bond isomerization in the periplasm is less efficient than in eukaryotes, and many eukaryotic proteins misfold into insoluble aggregates called inclusion bodies. Proteins larger than ~60 kDa are often poorly expressed. Additionally, E. coli produces endotoxins (lipopolysaccharide) that contaminate preparations and must be removed for therapeutic or certain cell-based applications.
For proteins that fold well in the bacterial cytosol, E. coli Protein Expression is the fastest route from gene to purified protein. Strains such as BL21(DE3) are engineered to lack the Lon and OmpT proteases, reducing degradation of recombinant products.
Eukaryotic systems for complex proteins
When post-translational modifications are essential for function, eukaryotic hosts are required.
Yeast (Saccharomyces cerevisiae and Pichia pastoris) offers a middle ground. P. pastoris grows to very high cell densities on methanol and performs core N-glycosylation, though the glycan structures are high-mannose type and differ from mammalian patterns. Yeast is well suited for secreted proteins and can produce gram-per-liter yields. Disulfide bond formation is handled properly in the secretory pathway.
Insect cells (typically Sf9 or High Five cells infected with recombinant baculovirus) provide more authentic processing, including complex glycosylation closer to mammalian patterns, though still not identical. They are the system of choice for large multi-subunit complexes and membrane proteins that fail in yeast or bacteria. Yields are typically lower than bacterial systems—milligrams per liter rather than tens of milligrams—and the workflow is slower, requiring virus amplification.
Mammalian cells (HEK293, CHO) produce proteins with authentic human post-translational modifications, including complex sialylated glycans. This is mandatory for therapeutic proteins and often desirable for studying human proteins in their native state. However, mammalian culture is expensive, slow, and yields are modest (micrograms to milligrams per liter). Transient transfection of HEK293 cells is the fastest mammalian approach, while stable CHO cell lines are used for large-scale production.
The following table summarizes key considerations:
| Feature | E. coli | Yeast (P. pastoris) | Insect (Sf9/baculovirus) | Mammalian (HEK293/CHO) |
|---|---|---|---|---|
| Doubling time | ~20 min | ~90 min | ~24 h | ~24 h |
| Typical yield | 10–100 mg/L | 10–500 mg/L | 1–10 mg/L | 0.1–5 mg/L |
| Glycosylation | None | High-mannose | Complex (paucimannose) | Complex, sialylated |
| Disulfide bonds | Poor (cytosol) | Good (secretory) | Good | Good |
| Cost | Low | Low–moderate | Moderate–high | High |
| Time to protein | 2–5 days | 1–2 weeks | 2–4 weeks | 2–8 weeks |
| Best for | Prokaryotic proteins, unmodified eukaryotic proteins | Secreted proteins, high-yield eukaryotic | Multi-subunit complexes, membrane proteins | Human therapeutic proteins, native modifications |
Designing the Expression Construct
The expression construct is the single most important determinant of success. A well-designed construct can make a difficult protein tractable; a poorly designed one can doom an easy protein.
Key regulatory elements
The promoter controls when and how strongly the gene is transcribed. The T7 promoter in pET vectors is the most common choice for E. coli. T7 RNA polymerase is provided in trans by the host chromosome (in strains like BL21(DE3)) and is itself under control of the lacUV5 promoter, which is induced by isopropyl β-D-1-thiogalactopyranoside (IPTG). This two-tiered system provides very high transcription rates upon induction. The arabinose-inducible araBAD promoter (pBAD vectors) offers tighter regulation and tunable expression, useful for toxic proteins. The rhamnose-inducible system provides similar benefits.
The ribosome binding site (RBS) in bacteria, or the Kozak sequence in eukaryotes, determines translation initiation efficiency. In E. coli, the Shine-Dalgarno sequence (AGGAGG) must be positioned 5–9 nucleotides upstream of the start codon. The strength of the RBS can be tuned to modulate expression level.
Codon optimization
The genetic code is degenerate, and organisms differ in their codon usage preferences. If your gene contains codons that are rare in the expression host, translation stalls, leading to truncated products and low yields. Codon optimization replaces rare codons with frequent ones without altering the amino acid sequence. For E. coli, genes from GC-rich organisms (e.g., Mycobacterium tuberculosis, human) often benefit substantially. Many commercial gene synthesis services include codon optimization as standard. Alternatively, E. coli strains such as Rosetta (BL21(DE3)pLysS) supply tRNAs for rare codons (AGA, AGG, AUA, CUA, GGA, CCC) and can rescue expression without gene resynthesis.
Fusion tags
Fusion tags are peptide or protein sequences appended to the target to facilitate purification, enhance solubility, or both. The choice of tag is a major decision.
Polyhistidine tags (His-tags) — Six to ten consecutive histidine residues coordinate divalent metal ions (Ni²⁺, Co²⁺) immobilized on a chromatography resin. His-tags are small (0.8–1.4 kDa), minimally immunogenic, and function under denaturing conditions, making them useful for purifying inclusion body proteins. They can be placed at either terminus. The mechanism of binding and optimization is covered in detail in the section on His Tag Protein Purification.
Glutathione S-transferase (GST) tag — A 26 kDa protein that binds glutathione immobilized on agarose. GST tags dramatically enhance solubility of many eukaryotic proteins and allow gentle elution with reduced glutathione (10–20 mM). The large size, however, can interfere with structure-function studies and must be removed for most applications.
Maltose-binding protein (MBP) tag — A 40 kDa protein that binds amylose resin and is eluted with maltose (10 mM). MBP is arguably the most effective solubility enhancer among common tags, often rescuing proteins that are otherwise completely insoluble. Its chaperone-like activity is attributed to its own folding pathway, which appears to nucleate folding of the fused passenger.
Small epitope tags (FLAG, HA, c-Myc) — Short peptides (8–12 residues) recognized by specific antibodies. They are useful for detection and gentle immunoaffinity purification but offer no solubility benefit and are expensive for preparative work.
Cleavage strategies to remove tags
Tags must be removed for most structural, biophysical, and therapeutic applications. The standard approach is to include a protease recognition site between the tag and the target protein.
TEV protease (tobacco etch virus) recognizes the seven-residue sequence ENLYFQ↓G and cleaves between Q and G. It is highly sequence-specific, active at 4 °C, and can be produced in-house as a His-tagged recombinant protein that is easily removed by passing the cleavage reaction over a Ni-NTA column. TEV leaves a single glycine (or serine) at the N-terminus of the target, which is usually acceptable.
Thrombin recognizes LVPR↓GS and is widely used but less specific than TEV, with occasional off-target cleavage. Factor Xa recognizes IEGR↓ and is also used, though it can cleave at secondary sites. PreScission protease is a GST-tagged version of human rhinovirus 3C protease that recognizes LEVLFQ↓GP and can be removed with glutathione resin.
For all protease cleavage strategies, the cleavage site must be accessible. A flexible linker (e.g., GGGGS) between the tag and the protease site often improves cleavage efficiency. After cleavage, a second affinity step removes both the protease and the cleaved tag, leaving the untagged target in the flow-through.
Cell Lysis and Solubilization
Efficient cell lysis is critical for maximizing yield and preserving protein integrity. The method must break the cell wall and membrane without denaturing the target protein or releasing excessive amounts of host proteases.
Mechanical vs. non-mechanical lysis
Sonication is the most common laboratory method. High-frequency sound waves (20 kHz) create cavitation bubbles that collapse and generate shear forces that disrupt membranes. A typical protocol for E. coli from 1 L culture: resuspend the cell pellet in 30–50 mL lysis buffer (50 mM Tris-HCl pH 8.0, 300 mM NaCl, 1 mM phenylmethylsulfonyl fluoride (PMSF) or a protease inhibitor cocktail, and optionally 1 mg/mL lysozyme and a few units of Benzonase nuclease to degrade nucleic acids). Sonicate on ice in 10-second pulses at 40–60% amplitude with 20-second rests, for a total of 3–5 minutes of active sonication. The lysate should become translucent. Over-sonication causes heating and protein denaturation; under-sonication leaves cells intact.
French press uses high pressure (20,000–30,000 psi) to force cells through a narrow orifice, causing shear. It is gentler on proteins than sonication, generates less heat, and is more reproducible, but requires specialized equipment. Microfluidizers operate on a similar principle and are scalable.
Enzymatic lysis — Lysozyme (from hen egg white) degrades the peptidoglycan layer of bacterial cell walls. It is often used in combination with mild detergent (0.1–1% Triton X-100) and freeze-thaw cycles. This approach is gentler than mechanical methods and is preferred for fragile proteins, but it is slower and less complete for high-density cultures.
Non-mechanical alternatives include osmotic shock, freeze-thaw cycling, and detergent-based lysis. These are generally less efficient for bacteria but are the methods of choice for mammalian cells, which lack cell walls and lyse readily in hypotonic buffers or with mild detergents.
Inclusion body handling
When a recombinant protein is expressed at high levels in E. coli, it often aggregates into dense, insoluble particles called inclusion bodies. These are visible by phase-contrast microscopy and can be recovered by low-speed centrifugation (12,000 × g for 15 min) after lysis. Inclusion bodies consist of partially folded or misfolded protein, often in a β-sheet-rich amyloid-like state, along with nucleic acids and membrane components.
Inclusion bodies are not necessarily a dead end. They offer advantages: the protein is protected from proteolysis, is highly enriched (often >50% pure), and can be obtained in large quantities. The challenge is refolding.
The standard workflow for inclusion body proteins:
- Wash the inclusion body pellet with buffer containing 2 M urea and 1–2% Triton X-100 to remove membrane contaminants.
- Solubilize in a strong chaotrope: 6–8 M guanidine hydrochloride (GuHCl) or 8 M urea, in 50 mM Tris pH 8.0, with 5–10 mM dithiothreitol (DTT) or β-mercaptoethanol to reduce disulfide bonds. Incubate at room temperature for 1–2 hours with stirring.
- Clarify by centrifugation (30,000 × g, 30 min) to remove insoluble debris.
- Refold by removing the denaturant. This is the critical step. Methods include:
- Dialysis against a refolding buffer (e.g., 50 mM Tris pH 8.0, 150 mM NaCl, 1 mM DTT, 0.5 M arginine) with gradual removal of denaturant over 12–24 hours.
- Rapid dilution — dropwise addition of the denatured protein into a large volume of refolding buffer (typically 10–50-fold excess) with gentle stirring.
- On-column refolding — bind the denatured protein to an affinity column (e.g., Ni-NTA for His-tagged proteins) in denaturing buffer, then gradually replace the buffer with refolding buffer while the protein is immobilized. This prevents aggregation by physically separating molecules.
Refolding yields are often poor (5–30%), and optimization of pH, ionic strength, redox conditions (for disulfide-containing proteins), and additives (arginine, glycerol, detergents) is usually required. If inclusion bodies are the only option, consider whether a different expression strategy—lower temperature, weaker promoter, or a solubility-enhancing fusion tag—might produce soluble protein instead.
Affinity Purification Techniques
Affinity chromatography exploits the specific, reversible interaction between a tag on the target protein and a ligand immobilized on a resin. This single step typically provides 50–100-fold purification, often yielding protein that is >80% pure.
IMAC: mechanism and optimization
Immobilized metal affinity chromatography (IMAC) is the most widely used affinity method, primarily for His-tagged proteins. The principle is the coordination of histidine imidazole side chains to transition metal ions (Ni²⁺, Co²⁺, Cu²⁺, Zn²⁺) chelated to the resin via nitrilotriacetic acid (NTA) or iminodiacetic acid (IDA). Ni-NTA is the standard choice, offering high capacity (5–10 mg His-tagged protein per mL resin) and good selectivity.
The detailed protocol for His Tagged Protein Purification follows a standard pattern:
- Equilibrate the column with binding buffer: 50 mM sodium phosphate pH 8.0, 300 mM NaCl, 10–20 mM imidazole. The imidazole competes with histidine side chains for metal coordination, reducing non-specific binding of host proteins that have surface-exposed histidines.
- Load the clarified lysate. Use a slow flow rate (0.5–1 mL/min for gravity columns) to allow binding. Collect the flow-through for analysis.
- Wash with binding buffer containing 20–50 mM imidazole to remove weakly bound contaminants.
- Elute with 250–500 mM imidazole. Collect fractions and analyze by SDS-PAGE.
Key optimization parameters:
- Imidazole concentration in load/wash: Too low (0–5 mM) increases background binding; too high (>30 mM) causes the target to flow through. The optimal concentration depends on the His-tag length and accessibility.
- pH: Histidine coordination requires the imidazole nitrogen to be deprotonated (pKa ~6.0). Binding is stronger at pH 8.0 than at pH 7.0.
- Salt: 300 mM NaCl reduces ionic interactions with contaminating nucleic acids and acidic proteins.
- Metal ion: Ni²⁺ has the highest capacity but also the highest background. Co²⁺ (TALON resin) gives higher purity at the cost of lower capacity.
Elution strategies and tag removal
Imidazole elution is standard but can co-elute contaminating proteins that bind weakly to the resin. Alternatives include:
- Low pH elution (e.g., 50 mM sodium acetate pH 4.5) — protonates the histidine imidazole, disrupting coordination. This can precipitate acid-sensitive proteins.
- EDTA elution — chelates the metal ion itself, releasing both target and any metal-bound contaminants. This strips the column and requires recharging with metal before reuse.
- On-column cleavage — if the His-tag is separated from the target by a protease site, the protease can be applied directly to the column. The target elutes while the tag and protease remain bound. This is elegant but requires careful optimization.
After elution, the His-tag is removed by protease cleavage as described in the construct design section. The cleavage reaction is then passed over a fresh Ni-NTA column; the cleaved His-tag, the His-tagged protease, and any uncleaved fusion protein bind, while the untagged target flows through. This subtractive step is the standard approach for Custom Protein Purification workflows.
Alternatives to IMAC
GST affinity — Glutathione-agarose binds GST-tagged proteins with high specificity. Elution with 10–20 mM reduced glutathione in 50 mM Tris pH 8.0 is gentle and preserves activity. GST resin is more expensive than Ni-NTA and has lower capacity, but the selectivity is often superior.
MBP affinity — Amylose resin binds MBP-tagged proteins; elution with 10 mM maltose. This is gentle and highly specific.
Immunoaffinity — Antibodies against the tag (e.g., anti-FLAG) or against the protein itself provide the highest selectivity but are expensive, have low capacity, and require harsh elution conditions (low pH or competing peptide).
Polishing and Analytical Characterization
Affinity purification rarely yields protein pure enough for structural studies or therapeutic use. Polishing steps remove residual contaminants, aggregated protein, and the cleaved tag.
Ion exchange and gel filtration
Ion exchange chromatography (IEX) separates proteins by surface charge. The choice of resin depends on the protein's isoelectric point (pI):
- Anion exchange (e.g., Q Sepharose, positively charged) binds negatively charged proteins at pH above their pI.
- Cation exchange (e.g., SP Sepharose, negatively charged) binds positively charged proteins at pH below their pI.
A typical protocol: dialyze the protein into a low-salt buffer (e.g., 20 mM Tris pH 8.0 for anion exchange), load onto the column, wash, and elute with a linear gradient of 0–500 mM NaCl. Proteins elute at characteristic salt concentrations based on their charge density. IEX also removes nucleic acids, which bind strongly to anion exchangers, and endotoxins.
Size exclusion chromatography (SEC), also called gel filtration, separates by hydrodynamic radius. It is the final polishing step, removing aggregates and exchanging the protein into the final storage buffer. A Superdex 200 column (for proteins up to ~600 kDa) or Superose 6 (for larger complexes) is typical. The protein is loaded in a small volume (0.5–5% of column volume) and eluted isocratically. Monodisperse protein elutes as a single symmetric peak; aggregates elute in the void volume. SEC is also the most reliable method for assessing the oligomeric state and homogeneity of the preparation.
Quality control: SDS-PAGE and mass spec
SDS-PAGE (sodium dodecyl sulfate-polyacrylamide gel electrophoresis) is the first-line assessment of purity. Samples are denatured, coated with negatively charged SDS, and separated by molecular weight. Coomassie Blue staining detects protein bands down to ~50 ng. A single band at the expected molecular weight indicates high purity, but the absence of visible contaminants does not guarantee homogeneity—minor species below the detection limit may still be present.
Western blotting confirms the identity of the protein using a specific antibody, either against the protein itself or against a tag. This is essential when the protein is expressed at low levels or when the SDS-PAGE band is ambiguous.
Mass spectrometry provides definitive identification. Intact protein mass analysis by electrospray ionization (ESI) or matrix-assisted laser desorption/ionization (MALDI) confirms the molecular weight and detects post-translational modifications. Peptide mass fingerprinting (trypsin digestion followed by MS/MS) confirms the primary sequence and can identify co-purifying contaminants.
Additional characterization methods include:
- UV absorbance — A280 measurement gives concentration (using the extinction coefficient calculated from the amino acid sequence). The A260/A280 ratio indicates nucleic acid contamination (>0.6 suggests significant contamination).
- Dynamic light scattering (DLS) — Assesses particle size distribution and detects aggregation.
- Size exclusion chromatography with multi-angle light scattering (SEC-MALS) — Provides absolute molecular weight of the native protein, confirming oligomeric state.
- Activity assays — The ultimate test of functionality, specific to each protein.
Optimizing Yield and Solubility
When expression fails—either low yield or insoluble protein—systematic optimization is required. The most powerful variables are temperature, inducer concentration, and the co-expression of folding helpers.
Lower temperature and slower induction
High expression rates overwhelm the protein folding machinery, leading to aggregation. Reducing the growth temperature after induction slows transcription and translation, giving the protein more time to fold. A common strategy:
- Grow cells at 37 °C to mid-log phase (OD₆₀₀ = 0.6–0.8).
- Cool the culture to 16–25 °C (15–30 min with shaking).
- Induce with a reduced IPTG concentration (0.1–0.5 mM instead of 1 mM).
- Incubate for 12–24 hours at the lower temperature.
For many proteins, this simple change converts an entirely insoluble expression into a predominantly soluble one. The trade-off is lower total yield—cells grow more slowly at lower temperatures—but the soluble yield is often higher because less protein is lost to inclusion bodies.
Co-expression with molecular chaperones
Molecular chaperones assist protein folding by binding exposed hydrophobic surfaces and preventing aggregation. Co-expression of chaperone systems can rescue difficult proteins.
The most commonly used systems:
- DnaK-DnaJ-GrpE (Hsp70 system) — Binds nascent polypeptide chains and facilitates folding.
- GroEL-GroES (Hsp60 system) — Provides an enclosed cage for folding of larger proteins.
- Trigger factor — Ribosome-associated chaperone that binds nascent chains co-translationally.
These are available in commercial plasmids (e.g., pKJE7, pGro7) with compatible origins of replication and antibiotic markers, allowing co-transformation with the expression plasmid. Induction of chaperone expression with arabinose or tetracycline precedes induction of the target protein.
Other additives that improve solubility:
- Ethanol (2–3%) or sorbitol (0.5 M) — Osmolytes that stabilize protein structure.
- Betaine (1 M) — Compatible solute that promotes folding.
- Glycerol (5–10%) — Stabilizes proteins and reduces aggregation.
- Arginine (0.5 M) — Suppresses aggregation during refolding and elution.
Troubleshooting insoluble proteins
If the protein remains insoluble despite optimization, consider:
- Is the protein inherently aggregation-prone? Predictions from sequence (e.g., high β-sheet content, exposed hydrophobicity) may indicate that the protein requires a eukaryotic chaperone system or a solubility-enhancing fusion partner.
- Try a different fusion tag. MBP is the most effective solubility enhancer; GST is also useful. The tag can be removed after purification.
- Express as a fusion with a small, well-folding protein such as SUMO (small ubiquitin-like modifier). SUMO fusions often improve folding and can be cleaved by SUMO protease (Ulp1), which leaves no residual amino acids.
- Secrete the protein to the periplasm in E. coli using a signal peptide (e.g., PelB, OmpA). The oxidizing environment of the periplasm supports disulfide bond formation.
- Consider a different expression host. If bacterial expression fails repeatedly, the protein may require eukaryotic folding machinery. Contract Recombinant Protein Expression services can test multiple systems in parallel.
Common Pitfalls and Troubleshooting
Even with careful planning, problems arise. Here are the most frequent failure modes and systematic approaches to resolving them.
Dealing with degradation
Proteolytic degradation manifests as multiple lower-molecular-weight bands on SDS-PAGE or a smeared pattern. Contributing factors include:
- Host proteases — Use protease-deficient strains (BL21(DE3)pLysS, Rosetta). Add protease inhibitors to the lysis buffer: PMSF (1 mM), leupeptin (1 µg/mL), pepstatin (1 µg/mL), and EDTA (1 mM, which inhibits metalloproteases).
- Cell lysis releases proteases — Work quickly at 4 °C. Keep the lysate on ice at all times.
- The protein is intrinsically unstable — The N-end rule states that proteins with bulky or basic N-terminal residues (Arg, Lys, Phe, Leu) are rapidly degraded. If the tag is removed, the exposed N-terminus may destabilize the protein. Consider an N-terminal stabilizing residue or a different cleavage strategy.
- Degradation during purification — Include protease inhibitors in all buffers. If degradation occurs after lysis, the problem is likely a specific contaminating protease; identify it by mass spectrometry and adjust inhibitors accordingly.
Avoiding co-purifying contaminants
The most common contaminants in affinity purification:
- Chaperones (DnaK, GroEL) — These bind to exposed hydrophobic surfaces of partially folded proteins. They co-purify with His-tagged proteins because they bind the target, not the resin. Improve folding (lower temperature, chaperone co-expression) or add ATP (5 mM) with MgCl₂ to the wash buffer to release chaperones.
- Nucleic acids — DNA and RNA bind to many proteins and to IMAC resins. Add Benzonase (25 U/mL) to the lysis buffer and include 300 mM NaCl in all buffers. If contamination persists, add a heparin or anion exchange step.
- Endotoxins — Lipopolysaccharide from E. coli is a major concern for cell-based assays. Remove by Triton X-114 phase separation, polymyxin B affinity, or anion exchange.
- Non-specific His-rich proteins — Endogenous E. coli proteins with surface-exposed histidines bind Ni-NTA. Increase the imidazole concentration in the wash buffer (up to 50–75 mM) or switch to Co²⁺ resin.
Low yield from the affinity column
If the target protein does not bind to the affinity resin:
- Is the tag accessible? The tag may be buried in the folded protein. Try moving it to the other terminus or adding a flexible linker.
- Is the tag present? Verify by Western blot with an anti-His antibody. If the tag is cleaved off by host proteases, use a protease-deficient strain or add inhibitors.
- Is the imidazole concentration too high? Reduce to 5–10 mM in the load buffer.
- Is the resin saturated? Check the binding capacity. For high-expression proteins, use more resin or load in multiple batches.
- Is the protein in the insoluble fraction? Check both supernatant and pellet after lysis. If the protein is in inclusion bodies, see the refolding section.
Aggregation during purification
Protein aggregation is indicated by precipitation, turbidity, or a high-molecular-weight smear on native PAGE. Strategies:
- Add glycerol (5–10%) to all buffers.
- Increase salt to 300–500 mM NaCl to reduce hydrophobic interactions.
- Add a reducing agent (1–5 mM DTT or TCEP) if the protein contains cysteines.
- Work at 4 °C and minimize handling time.
- Avoid freeze-thaw cycles — aliquot and snap-freeze in liquid nitrogen, store at −80 °C.
Summary and Best Practices
The successful production of a pure, functional protein is an iterative process that rewards careful planning and systematic troubleshooting. The following best practices will serve you well:
- Design the construct with purification in mind. Choose the tag, protease site, and expression system based on the protein's properties and intended use. A few hours of additional design work can save weeks of troubleshooting.
- Test expression on a small scale first. Before committing to a 4 L culture, test 5–50 mL cultures to assess expression level, solubility, and stability. This is where Custom Recombinant Protein Expression services can provide rapid screening.
- Keep a detailed notebook. Record every variable: strain, plasmid, induction conditions, buffer compositions, and results. This documentation is invaluable when problems arise.
- Analyze every fraction. Run SDS-PAGE on the lysate, flow-through, washes, and elution fractions. This tells you where the protein is being lost and guides optimization.
- Optimize one variable at a time. Change temperature, inducer concentration, or buffer composition individually. Changing multiple variables simultaneously makes it impossible to identify the cause of improvement or failure.
- Be patient and systematic. Protein purification is rarely successful on the first attempt. Each failure provides information that guides the next iteration.
Frequently Asked Questions
What is the best expression system for my protein?
There is no universal answer. Start with E. coli if your protein is prokaryotic, small (<60 kDa), and requires no glycosylation. If the protein is eukaryotic, try E. coli first with solubility-enhancing tags; if it fails, move to yeast (P. pastoris) for secreted proteins, then insect or mammalian cells for proteins requiring complex post-translational modifications. Consider the Recombinant Protein Expression System options available in your lab or through commercial services.
How do I increase protein solubility during expression?
The most effective strategies, in order: (1) reduce the growth temperature after induction to 16–25 °C; (2) reduce the IPTG concentration to 0.1–0.5 mM; (3) use a solubility-enhancing fusion tag (MBP, GST, SUMO); (4) co-express molecular chaperones (GroEL-GroES, DnaK-DnaJ-GrpE); (5) add osmolytes (sorbitol, betaine) to the culture medium.
Why is my protein not binding to the affinity column?
Check the tag is present and accessible (Western blot), reduce the imidazole concentration in the load buffer, verify the resin is charged with the correct metal ion, and confirm the protein is in the soluble fraction. If the protein is in inclusion bodies, it will not bind unless you purify under denaturing conditions (6 M GuHCl or 8 M urea in all buffers).
What is the typical yield of purified protein from E. coli?
For a well-expressed, soluble protein, expect 5–50 mg of purified protein per liter of shake-flask culture. High-density fermentation can increase this to 100–500 mg/L. If your yield is below 1 mg/L, the expression or purification protocol needs optimization.
How do I remove the His-tag after purification?
Incorporate a protease cleavage site (TEV, thrombin, Factor Xa) between the His-tag and the target. After affinity purification, add the protease (typically 1:100 protease:protein by weight) and incubate at 4 °C for 12–16 hours. Pass the cleavage reaction over a fresh Ni-NTA column; the cleaved His-tag, the His-tagged protease, and uncleaved protein bind, while the untagged target flows through.
What are inclusion bodies and how do I handle them?
Inclusion bodies are insoluble aggregates of misfolded protein formed during high-level expression in E. coli. They are recovered by centrifugation, washed, solubilized in 6–8 M GuHCl or 8 M urea, and refolded by gradual removal of the denaturant. Refolding yields are often low (5–30%), so it is usually better to prevent inclusion body formation by optimizing expression conditions.
How do I check the purity of my purified protein?
Run SDS-PAGE with Coomassie staining to assess purity. For higher sensitivity, use silver staining or Western blot. Confirm identity by mass spectrometry (intact mass or peptide fingerprinting). Assess native state and oligomeric homogeneity by size exclusion chromatography or dynamic light scattering. For functional proteins, always perform an activity assay.
Key Takeaways
- Protein expression and purification is a multi-stage pipeline; decisions made at the construct design stage propagate through every subsequent step.
- Choose the expression system based on the protein's origin, size, complexity, and required post-translational modifications; E. coli is the fastest and cheapest starting point.
- Fusion tags (His, GST, MBP) enable affinity purification and can enhance solubility; always include a protease cleavage site for tag removal.
- IMAC with Ni-NTA is the workhorse of affinity purification; optimize imidazole concentration, pH, and salt to maximize selectivity.
- Polishing by ion exchange and size exclusion chromatography is essential for high-purity applications; always characterize the final product by SDS-PAGE and mass spectrometry.
- Low temperature, reduced inducer concentration, and chaperone co-expression are the most effective levers for improving protein solubility.
- Troubleshooting requires systematic variation of one parameter at a time, careful analysis of every fraction, and meticulous record-keeping.
Further Reading
- Wingfield PT. Overview of the purification of recombinant proteins. Current protocols in protein science. 2015. PubMed 25829302
- Růčková E, Müller P, Vojtěšek B. [Protein expression and purification]. Klinicka onkologie : casopis Ceske a Slovenske onkologicke spolecnosti. 2014. PubMed 24945544
- Kim Y et al. High-throughput protein purification and quality assessment for crystallization. Methods (San Diego, Calif.). 2011. PubMed 21907284
- Young CL, Britton ZT, Robinson AS. Recombinant protein expression and purification: a comprehensive review of affinity tags and microbial applications. Biotechnology journal. 2012. PubMed 22442034
- Lin Z et al. Cleavable self-aggregating tags (cSAT) for protein expression and purification. Methods in molecular biology (Clifton, N.J.). 2015. PubMed 25447859
- Lesley SA. Parallel methods for expression and purification. Methods in enzymology. 2009. PubMed 1989220163041-X)