How to Calculate the Isoelectric Point (pI) of an Amino Acid or Protein
By Dr. Zubair Khalid, DVM, MS, PhD ·

The isoelectric point (pI) is the pH at which a molecule carries no net charge. For a free amino acid, that means the population is dominated by the neutral zwitterion, with the positively charged alpha-amino group and the negatively charged alpha-carboxyl group balanced against each other [2]. For a protein or peptide, the pI is the pH at which the sum of all positive and negative charges on the molecule equals zero, so the protein stops moving in an electric field [7].
You will meet this number in several places. In a 2-D gel, a protein migrates until it reaches the pH in the strip that matches its pI, which is the basis of isoelectric focusing [8]. In ion exchange chromatography, the choice between a cation and an anion exchanger depends on whether your protein is above or below its pI at the working pH. In formulation and purification, proteins are usually least soluble at their pI because the net charge is zero and molecules aggregate more readily [2]. Getting the calculation right, and knowing which pKa set produced it, saves time at the bench.
Quick Answer
- The pI is the pH at which net charge is zero. For a protein, it is the pH where electrophoretic mobility stops [7].
- For a free amino acid, the pI is the average of the two pKa values that flank the neutral zwitterion [2].
- For the 13 amino acids with neutral side chains, pI = (pKa1 + pKa2)/2, where pKa1 is the alpha-carboxyl and pKa2 is the alpha-ammonium group [2].
- For acidic side chains (Asp, Glu, Cys, Tyr), average the two lowest pKa values. For basic side chains (Lys, Arg, His), average the two highest [2].
- For peptides and proteins, use the charge-balance equation and solve for the pH where total charge is zero. The averaging rule does not apply because many groups ionize at once [7].
- The answer depends on the pKa set you choose [7]. In the peptide example below, switching tables shifts the calculated pI by about 0.8 pH units.
What the Isoelectric Point Actually Means
Every ionizable group on a molecule has a pKa, the pH at which it is half protonated and half deprotonated. The Henderson-Hasselbalch equation describes this balance:
$$ \mathrm{pH} = \mathrm{p}K_a + \log\frac{[\mathrm{A^-}]}{[\mathrm{HA}]} $$
Here [A⁻] is the concentration of the deprotonated (conjugate base) form, [HA] is the protonated (acid) form, and pKa is the negative log of the acid dissociation constant for that group. When pH equals pKa, the two forms are present in equal amounts [2].
Amino acids carry at least two ionizable groups: the alpha-carboxyl group, which is acidic, and the alpha-amino group, which is basic. Seven side chains are also ionizable: Asp, Glu, Cys, and Tyr are acidic, while His, Lys, and Arg are basic [7]. At low pH, the acidic groups are protonated and neutral, and the basic groups are protonated and positive, so the molecule has a net positive charge. At high pH, the acidic groups lose their protons and become negative, and the basic groups lose theirs and become neutral, so the net charge is negative. Somewhere in between, the positive and negative charges cancel.
That crossover point is the pI. NLM MeSH defines it as the pH at which dipolar ions are at a maximum [8]. The definition matters because it is a statement about net charge, not about the absence of charge. A protein at its pI can still carry many positive and negative charges internally; they simply sum to zero.
How to Calculate the pI of a Free Amino Acid
The averaging rule comes directly from the shape of the titration curve. For a simple amino acid with a neutral side chain, there are two ionization steps. The zwitterion exists between them, so the pH at which the zwitterion dominates sits halfway between the two pKa values:
$$ \mathrm{pI} = \frac{\mathrm{p}K_{a1} + \mathrm{p}K_{a2}}{2} $$
For glycine, pKa1 is 2.34 and pKa2 is 9.60, giving a pI of 5.97 [1]. The logic extends to amino acids with a third ionizable group. If the side chain is acidic, it adds an extra negative charge, so the neutral form lies between the two lowest pKa values. If the side chain is basic, it adds an extra positive charge, so the neutral form lies between the two highest pKa values [2].
The rule works because only two ionization steps flank the zwitterion. Once you move to a peptide or protein, dozens of groups ionize across a wide pH range, and no single pair of pKa values defines the neutral state. That is where the charge-balance approach takes over.
How to Calculate the pI of a Peptide or Protein
For a protein, the net charge at a given pH is the sum of contributions from every ionizable group. The standard model is:
$$ Q(\mathrm{pH}) = \sum_i \frac{n_i}{1 + 10^{(\mathrm{pH} - \mathrm{p}K_{a,i})}} - \sum_j \frac{n_j}{1 + 10^{(\mathrm{p}K_{a,j} - \mathrm{pH})}} $$
The first sum runs over basic groups (N-terminus, His, Lys, Arg), where nᵢ is the count of that group and pKa,i is its pKa. The second sum runs over acidic groups (C-terminus, Asp, Glu, Cys, Tyr), where nⱼ is the count and pKa,j is its pKa. The pI is the pH at which Q equals zero [7].
There is no closed-form solution for that pH, so you solve it numerically. Bisection or Brent's method works well because Q decreases monotonically as pH rises. You bracket a pH range, evaluate Q at both ends, and narrow the interval until Q is close enough to zero. A tolerance of 0.001 pH units is more than sufficient given the underlying uncertainty in pKa values.
The pKa values themselves are empirical. They vary with temperature, ionic strength, and the specific experimental conditions under which they were measured. Kozlowski (2016) compared 15 published pKa sets and found that the choice of set changes the predicted pI [7]. Expasy's Compute pI/Mw uses pK values from Bjellqvist et al., derived from polypeptide migration in immobilized pH gradient gels containing urea [3]. EMBOSS uses its own Epk.dat table [6]. Biopython implements the Bjellqvist method [9]. If you need a quick number for a protein you are purifying, the Protein Properties Calculator on this site will calculate a pI from your sequence.
Worked Example
Part A: Free amino acids
Glycine. The pKa values are 2.34 for the alpha-carboxyl group and 9.60 for the alpha-ammonium group [1]. The zwitterion forms between these two steps, so:
$$ \mathrm{pI} = \frac{2.34 + 9.60}{2} = \frac{11.94}{2} = 5.97 $$
Aspartate. The pKa values are 1.88 (alpha-COOH), 3.65 (side-chain COOH), and 9.60 (alpha-NH₃⁺) [1]. The neutral zwitterion lies between the two lowest steps:
$$ \mathrm{pI} = \frac{1.88 + 3.65}{2} = \frac{5.53}{2} = 2.765 \approx 2.77 $$
Lysine. The pKa values are 2.18 (alpha-COOH), 8.95 (alpha-NH₃⁺), and 10.53 (side-chain NH₃⁺) [1]. The neutral zwitterion lies between the two highest steps:
$$ \mathrm{pI} = \frac{8.95 + 10.53}{2} = \frac{19.48}{2} = 9.74 $$
You can check each result by solving Q = 0 with the charge-balance equation. Using scipy.optimize.brentq with the OpenStax pKa values gives 5.9700 for glycine, 2.7650 for aspartate, and 9.7400 for lysine. At pH 7.0, the net charges are -0.002, -1.002, and +0.989 respectively. Glycine is nearly neutral at physiological pH, aspartate is negative, and lysine is positive, which matches what the pI values tell you.
Part B: A peptide
Human angiotensin II is the octapeptide DRVYIHPF, corresponding to residues 25 to 32 of angiotensinogen (UniProt P01019) [10]. Its ionizable groups are the N-terminal amine, Arg, His, Asp, Tyr, and the C-terminal carboxyl.
Using Biopython 1.88:
from Bio.SeqUtils.ProtParam import ProteinAnalysis
pa = ProteinAnalysis('DRVYIHPF')
pa.isoelectric_point() # returns 6.7436
pa.charge_at_pH(7.0) # returns -0.1526
pa.charge_at_pH(7.4) # returns -0.4080
An independent implementation with the same Bjellqvist pK values (N-terminus 7.5, Arg 12.0, His 5.98, C-terminus 3.55, Asp 4.05, Tyr 10.0) and Brent's method gives the same pI of 6.7436. The net charge curve runs from +2.697 at pH 3, through +1.037 at pH 5, -0.153 at pH 7, -0.408 at pH 7.4, -1.060 at pH 9, and -2.000 at pH 11.
Now switch to the EMBOSS Epk.dat pKa values (N-terminus 8.6, Arg 12.5, His 6.5, C-terminus 3.6, Asp 3.9, Tyr 10.1). The same peptide gives a pI of 7.543 and a charge of +0.216 at pH 7.0. That is a shift of about 0.8 pH units, driven mainly by the N-terminal and histidine pKa values, not by anything in the sequence. The lesson is simple: always report which pKa set produced your number.
Reading and Using the Result
A pI above 7 means the protein carries a net positive charge at neutral pH. A pI below 7 means it carries a net negative charge. That single fact determines your choice of ion exchange resin. If your protein has a pI of 6.7 and you run your column at pH 7.4, the protein is negative and will bind to an anion exchanger. Run the same column at pH 6.0 and the protein is positive and will bind to a cation exchanger.
In electrophoresis, the same logic applies. Proteins in a buffer above their pI are negatively charged and migrate toward the positive electrode. Below their pI, they are positive and move toward the negative electrode [2]. In isoelectric focusing, a pH gradient is established in a gel, and each protein migrates until it reaches the position where the pH equals its pI, at which point it stops [8]. This is why IEF can resolve proteins that differ by a single charge.
Solubility follows the same pattern. At the pI, net charge is zero, electrostatic repulsion between molecules is minimal, and the protein is most likely to aggregate or precipitate. Moving the pH away from the pI in either direction increases the net charge and usually improves solubility [2]. This is the principle behind isoelectric precipitation, a common early step in protein purification.
Common Mistakes
- Applying the averaging rule to a peptide or protein. The rule works for free amino acids because exactly two ionization steps flank the zwitterion. A peptide has many ionizable groups, and no single pair of pKa values defines its neutral state. Use charge-balance root finding instead [7].
- Reporting a pI without stating the pKa set. Bjellqvist, EMBOSS, and other tables give different answers for the same sequence [7]. The 0.8-unit gap for angiotensin II is entirely due to the pKa table, not the calculation method.
- Treating the calculated pI as an exact experimental value. In Kozlowski's 25% test set, the RMSD between predicted and measured pI was 0.874 pH units for the best-performing method [7]. A calculated pI of 6.74 does not mean the protein will focus at exactly 6.74.
- Assuming post-translational modifications are included. Expasy Compute pI/Mw and ProtParam do not account for phosphorylation, glycosylation, or other modifications [3][4]. Glycosylation in particular shifts a protein's position in both the pI and mass dimensions of a 2-D gel [3].
- Forgetting that disulfide bonds remove charge. When two cysteines form a disulfide, neither contributes a charge. A sequence with many cysteines will have a different effective pI under oxidizing conditions than the raw sequence suggests [7].
- Confusing the anode with the negative electrode. In electrophoresis, the anode is positive. Proteins below their pI are positive and migrate toward the cathode (negative electrode); proteins above their pI are negative and migrate toward the anode (positive electrode) [2].
Limitations
The charge-balance model assumes that each ionizable group behaves independently, with a pKa that does not change when neighboring groups ionize. Real proteins do not work that way. Solvent exposure, hydrogen bonding, and electrostatic interactions between nearby charges shift pKa values away from their model compound values [7]. EMBOSS states plainly that its calculation assumes no electrostatic interactions change the propensity for ionization [6].
The Bjellqvist pK values were derived under denaturing conditions, using urea at high concentration [3]. They predict the focusing positions of unfolded polypeptides on 2-D gels well, but they are less reliable for native folded proteins, where buried charges and salt bridges alter ionization behavior.
Expasy notes that pI prediction for highly basic proteins is not well studied, and that poor buffer capacity increases prediction error, making pI predictions for small proteins problematic [3]. The Bjellqvist pK values used here are those implemented in Biopython and documented by Expasy [5].
Textbook pKa values also differ slightly between sources. The OpenStax histidine side-chain pKa is 6.00 [1], while other tables place it near 6.0 to 6.1. These small differences matter less than the choice of pKa set overall, but they are worth noting when you compare results across tools.
Frequently Asked Questions
What is the isoelectric point in simple terms?
The isoelectric point is the pH at which a molecule has no net charge. For an amino acid, it is the pH where the neutral zwitterion dominates. For a protein, it is the pH where the total positive charge from basic groups exactly balances the total negative charge from acidic groups [7].
How do I find the isoelectric point of an amino acid?
Look up the pKa values for that amino acid. If the side chain is neutral, average the alpha-carboxyl and alpha-ammonium pKa values. If the side chain is acidic, average the two lowest pKa values. If the side chain is basic, average the two highest [2]. For glycine, that gives (2.34 + 9.60)/2 = 5.97 [1].
What is the isoelectric point formula for a protein?
There is no simple formula. You sum the charge contributions from every ionizable group using the Henderson-Hasselbalch equation, then find the pH where that sum equals zero [7]. The equation is Q(pH) = sum of nᵢ/(1 + 10^(pH - pKa,i)) for basic groups minus sum of nⱼ/(1 + 10^(pKa,j - pH)) for acidic groups. Solve numerically.
Why do different protein pI calculators give different answers?
They use different pKa tables. Expasy uses Bjellqvist values, EMBOSS uses its own Epk.dat values, and other tools use still other sets [3][6]. Kozlowski (2016) compared 15 published sets and found that the choice of set changes the predicted pI [7]. For angiotensin II, the Bjellqvist and EMBOSS sets differ by about 0.8 pH units.
How is isoelectric focusing related to pI?
Isoelectric focusing is the experimental technique that measures pI. A pH gradient is established in a gel, and proteins migrate through it until they reach the pH that matches their pI, at which point their net charge is zero and they stop moving [8]. The pI you calculate from sequence is a prediction of where that stopping point will be.
References
- OpenStax Organic Chemistry 26.1: Structures of Amino Acids (Table 26.1, pKa and pI values)
- OpenStax Organic Chemistry 26.2: Amino Acids and the Henderson-Hasselbalch Equation: Isoelectric Points
- Expasy Compute pI/Mw documentation (SIB)
- Expasy ProtParam documentation (SIB)
- Bjellqvist B, et al. The focusing positions of polypeptides in immobilized pH gradients can be predicted from their amino acid sequences. Electrophoresis. 1993;14(10):1023-1031
- EMBOSS iep documentation (isoelectric point program and Epk.dat pK table)
- Kozlowski LP. IPC: Isoelectric Point Calculator. Biol Direct. 2016;11:55
- NLM MeSH: Isoelectric Focusing (D007525)
- Biopython API documentation: Bio.SeqUtils.IsoelectricPoint
- UniProtKB P01019, Angiotensinogen (human), peptide feature Angiotensin-2
Related Articles
- Isoelectric Focusing (IEF) for Protein Analysis
- Ion Exchange Chromatography for Protein Purification
- Amino Acid Structure
- How to Prepare Buffer Solutions: A Step-by-Step Guide
- Amino Acid Properties: Side Chains, Charge, Polarity and What Substitutions Do
- How to Calculate Protein Concentration From A280 Using the Extinction Coefficient