# Amino Acid Properties: Side Chains, Charge, Polarity and What Substitutions Do

Proteins are built from a repertoire of 20 standard amino acids, and only L amino acids are constituents of proteins [1]. Every one of those 20 shares the same backbone: an alpha carbon bonded to an amino group, a carboxyl group, a hydrogen, and a side chain. The backbone is what links residues into a chain. The side chain, or R group, is what makes each amino acid different, and it is the physical chemistry of those side chains that decides whether a protein folds, where it binds, and which residues can be swapped without breaking it.

You will meet the properties of amino acids constantly. A mutagenesis experiment asks whether a change at one position matters. A multiple sequence alignment shows a column where every sequence carries lysine or arginine. A docking run flags a buried aspartate. A clinical variant report says a missense change is "probably damaging." In each case the reasoning runs through side chain size, charge, polarity and the substitution scores that summarize how often one residue replaces another in real protein families.

## Quick Answer

- The 20 standard amino acids differ only in their side chains, which vary in size, shape, polarity and charge [1].
- At neutral pH, free amino acids exist predominantly as dipolar ions (zwitterions), with the amino group protonated and the carboxyl group deprotonated [1].
- Side chains fall into broad groups: nonpolar aliphatic (G, A, V, L, I, M, P), aromatic (F, Y, W), polar uncharged (S, T, C, N, Q), acidic (D, E) and basic (K, R, H). This five-group scheme is a common convention, not a single cited standard.
- Whether a side chain is charged depends on its pKa relative to the local pH, and measured pKa values in folded proteins shift with environment, with standard deviations around 1 pH unit [4].
- Substitution matrices such as BLOSUM62 assign a score to every possible exchange of one amino acid for another, and a positive score means the pair is seen in conserved protein blocks more often than chance [6][7].
- A conservative substitution keeps a residue in the same property group; a radical substitution changes charge or breaks structure, and the score usually reflects that.

## What Side Chains Actually Do

The backbone of a polypeptide is chemically repetitive. Its hydrogen bond donors and acceptors are the same at almost every position (proline lacks the backbone NH [8]), which is why secondary structure (alpha helices and beta sheets) forms from backbone interactions alone. The side chains project outward from that scaffold, and their chemistry determines what the chain can do next.

Size matters first. Glycine has a single hydrogen atom as its side chain and, with two hydrogens on the alpha carbon, is the only achiral standard amino acid [1]. That tiny side chain lets glycine fit into all structures, and it is well suited to reverse turns [8]. At the other end, tryptophan carries an indole ring, and proline's ring structure makes it more conformationally restricted than any other amino acid, markedly influencing protein architecture [2].

Polarity and charge matter next. The larger aliphatic side chains are hydrophobic, meaning they tend to cluster together instead of contacting water [2]. That clustering is a major driving force in folding. Charged side chains do the opposite: they prefer water and often sit on the protein surface, though buried charges exist and are frequently functional.

A few side chains have special chemistry. Cysteine resembles serine but carries a sulfhydryl (thiol, SH) group, and pairs of sulfhydryls can form disulfide bonds that stabilize some proteins [1]. Serine, threonine and tyrosine contain hydroxyl (OH) groups attached to a hydrophobic side chain, which makes them good hydrogen bond partners and common sites of phosphorylation [1]. Asparagine and glutamine contain a terminal carboxamide, while aspartic acid and glutamic acid have a carboxylic acid in its place, a swap of the amide NH2 for an OH that turns a polar residue into a charged one [1].

The aromatic set is phenylalanine, tyrosine and tryptophan. Phenylalanine contains a phenyl ring in place of one hydrogen of alanine [1]. Tyrosine adds a hydroxyl to that ring, which is why it behaves as both aromatic and polar.

## Charge, pKa and the pH Problem

An amino acid side chain is charged when the surrounding pH is on the ionized side of its pKa. The relationship is the Henderson-Hasselbalch equation:

$$
\mathrm{pH} = \mathrm{p}K_a + \log_{10}\left(\frac{[\mathrm{A^-}]}{[\mathrm{HA}]}\right)
$$

Here $\mathrm{p}K_a$ is the negative log of the acid dissociation constant for the group, $[\mathrm{A^-}]$ is the concentration of the deprotonated form, and $[\mathrm{HA}]$ is the concentration of the protonated form. When pH equals pKa, the two forms are equal. One pH unit above pKa, the deprotonated form dominates roughly ten to one.

Textbook pKa values are measured on free amino acids or small model compounds. Inside a folded protein, the local environment shifts them. Grimsley, Scholtz and Pace compiled 541 measured values from 78 proteins and reported averages with substantial spread: Asp 3.5 +/- 1.2 (n = 139), Glu 4.2 +/- 0.9 (n = 153), His 6.6 +/- 1.0 (n = 131), Cys 6.8 +/- 2.7 (n = 25), Tyr 10.3 +/- 1.2 (n = 20), Lys 10.5 +/- 1.1 (n = 35), the C-terminus 3.3 +/- 0.8 (n = 22) and the N-terminus 7.7 +/- 0.5 (n = 16) [4]. Those standard deviations of about 1 pH unit are the point: an individual group in an unusual environment can sit well away from the average.

Arginine is a special case. The guanidinium group has an intrinsic pKa of 13.8 +/- 0.1 by potentiometry and NMR, substantially higher than the value of about 12 often cited in textbooks [5]. Because of that high pKa, arginine side chains are predominantly charged even at pH 10 and remain protonated under physiological conditions even when buried in a hydrophobic environment [5].

At physiological pH, Asp, Glu, Lys and Arg are charged and His is partly charged [1]. Histidine has an imidazole side chain with a pKa near 6, so it can be uncharged or positively charged near neutral pH depending on its local environment, and it is often found in enzyme active sites where that switchability is useful [1].

## Reading an Amino Acid Properties Chart

A properties chart, such as the site's [Amino Acid Chart](/tools/amino-acid-chart), groups the 20 residues by side chain chemistry. The five-group scheme used most often in teaching is:

| Group | Residues | Behavior at neutral pH |
|---|---|---|
| Nonpolar aliphatic | G, A, V, L, I, M, P | Hydrophobic, cluster away from water |
| Aromatic | F, Y, W | Largely hydrophobic, Y has a polar OH |
| Polar uncharged | S, T, C, N, Q | Hydrogen bond donors and acceptors |
| Acidic | D, E | Negatively charged |
| Basic | K, R, H | Positively charged (H partly) |

This grouping is a convention, not a fixed rule. Textbooks disagree at the edges: Gly, Cys, Tyr, Trp, His and Met move between "nonpolar," "polar" and "aromatic" depending on the book, and Berg uses four groups instead of five. Treat the chart as a starting point for reasoning, not a verdict.

Two residues deserve a note because their behavior is not captured by polarity alone. Proline lacks a backbone NH group and tends to disrupt both alpha helices and beta strands, while glycine fits into all structures and is well suited to reverse turns [8]. By far the most common cis peptide bonds in proteins are X-Pro linkages, which is a direct consequence of proline's ring [9]. Isoleucine and threonine each contain a second chiral center in the side chain [2].

## Substitution Matrices: Turning Properties into Scores

A substitution matrix assigns a score to every possible exchange of one amino acid for another and is used to measure similarity when aligning protein sequences [6]. The BLOSUM family was derived by Henikoff and Henikoff from about 2000 blocks of aligned sequence segments characterizing more than 500 groups of related proteins, and it improved alignments and searches over the older Dayhoff-model matrices [6].

The NCBI BLOSUM62 file describes itself as a "BLOSUM Clustered Scoring Matrix in 1/2 Bit Units" built from Blocks Database 5.0 with a cluster percentage of 62% or more, with entropy 0.6979 and expected score -0.5209 [7]. The scores are log-odds values: a positive score means the pair is observed in conserved blocks more often than expected by chance, and a negative score means it is observed less often.

The scale runs from identity scores of 4 (Ala, Ile, Leu, Ser, Val) up to 11 (Trp), with Cys at 9 and Pro at 7. Off-diagonal scores are mostly negative: only 21 of the 190 distinct off-diagonal pairs in BLOSUM62 have a positive score, and the highest are Ile-Val 3 and Phe-Tyr 3.

Common conservative substitutions score well: Ile-Leu 2, Leu-Met 2, Asp-Glu 2, Lys-Arg 2, Glu-Gln 2, Tyr-Trp 2, His-Tyr 2, Ser-Thr 1, Asn-Asp 1, Lys-Glu 1. Radical substitutions score poorly: Leu-Asp -4, Gly-Ile -4, Gly-Leu -4, Cys-Glu -4, Phe-Pro -4, Trp-Pro -4, Asp-Trp -4, Asn-Trp -4, Ile-Asp -3, Pro-Leu -3.

Some small swaps still score zero or negative: Ala-Gly 0, Cys-Ser -1, Trp-Gly -2, Trp-Cys -2, Arg-Asp -2. That last one is a charge reversal, and the score reflects how rarely it survives in a conserved family.

## Worked Example

Score substitutions with BLOSUM62. In Python:

```python
from Bio.Align import substitution_matrices
B = substitution_matrices.load('BLOSUM62')
print(B['I']['V'])  # returns 3.0
```

Biopython's built-in BLOSUM62 matches the NCBI file at all 400 standard amino acid pairs. The NCBI file is at https://ftp.ncbi.nih.gov/blast/matrices/BLOSUM62.

Now read a set of substitutions against that matrix.

Conservative (same property group): Ile to Val +3, Phe to Tyr +3, Leu to Ile +2, Asp to Glu +2, Lys to Arg +2, Glu to Gln +2, His to Tyr +2, Ser to Thr +1, Asn to Asp +1.

Neutral or mildly unfavorable: Ala to Gly 0, Cys to Ser -1, Arg to Asp -2 (charge reversal), Trp to Gly -2.

Radical (nonpolar to charged, or breaking structure): Ile to Asp -3, Pro to Leu -3, Leu to Asp -4, Gly to Ile -4, Phe to Pro -4.

Identities for reference: Trp 11, Cys 9, Pro 7, Gly 6, Ala/Ile/Leu/Ser/Val 4.

Reading the numbers: a positive score means the pair is seen in conserved protein blocks more often than chance, so a Lys to Arg change in a sequence alignment is rewarded (+2) while Leu to Asp is penalized (-4). A three-residue example: aligning LKD with IRE scores 2 + 2 + 2 = 6, while aligning LKD with DGL scores -4 + -2 + -4 = -10.

Note how property and score can disagree. Cys to Ser looks conservative by shape but scores -1, probably because a cysteine in a disulfide is rarely replaceable. The score comes from real alignments, not from a chemistry textbook.

## Common Mistakes

- Treating the five-group chart as settled fact. Gly, Cys, Tyr, Trp, His and Met are grouped differently across textbooks, and Berg uses four groups. Use the chart as a reasoning aid and check the specific source you are citing.
- Assuming textbook pKa values apply inside a protein. Measured values in folded proteins have standard deviations around 1 pH unit, and 2.7 for cysteine [4]. A buried histidine may not behave like a histidine at pH 7.
- Using the textbook arginine pKa of about 12. The intrinsic value is 13.8 +/- 0.1, and arginine stays charged even when buried [5].
- Reading a positive BLOSUM score as "this substitution is safe." A positive score means the pair appears in conserved blocks more often than chance [6]. It says nothing about whether a specific change preserves function in your protein.
- Reading a negative score as "this substitution is impossible." Negative scores are the norm: only 21 of 190 off-diagonal pairs are positive. Many functional proteins carry substitutions the matrix penalizes.
- Confusing a conservative substitution with a silent one. Conservative means same property group. Silent means the codon change does not alter the amino acid at all.
- Assuming one-letter codes are always interchangeable with three-letter codes. IUPAC-IUB recommends that one-letter symbols be restricted to the comparison of long sequences [3]. Many one-letter symbols are the first letter of the amino acid's name, and the others were assigned by convention [2].

## Limitations

The property groupings are pedagogical. There is no single authoritative partition of the 20 amino acids into polarity classes, and different textbooks draw the lines differently. When a claim depends on grouping, cite the specific scheme.

Measured pKa values are averages over heterogeneous environments. The Grimsley averages are not intrinsic model-compound values, and the spread is the real information: a single side chain's pKa can sit well away from the mean [4]. Arginine is not in that table at all, and the 13.8 value comes from a different method and differs from the value of about 12 often used in textbooks [5].

The log-odds definition of BLOSUM scores (score = 2 log2 of observed over expected pair frequency, in half-bit units) and the interpretation "positive means more often than chance" come from the Henikoff paper's methods [6]. The matrix header and the abstract support the practical use, but the derivation details are worth checking in the original if you plan to modify or re-derive a matrix.

BLOSUM62 is a general-purpose default. It was built to improve sequence alignments and database searches [6], not for scoring the effect of a specific missense variant in a specific protein. For clinical interpretation, pair it with structural data, conservation scores and functional assays.

## Frequently Asked Questions

### What is the difference between nonpolar amino acids and polar amino acids?

Nonpolar amino acids have side chains that do not form favorable interactions with water, so they tend to cluster in the protein interior. Polar amino acids carry groups that can hydrogen bond with water or with other residues. The larger aliphatic side chains are the classic nonpolar set [2], and in the five-group scheme above, serine, threonine, cysteine, asparagine and glutamine form the polar uncharged set.

### Which amino acids are charged at physiological pH?

Aspartate, glutamate, lysine and arginine are charged at physiological pH, and histidine is partly charged [1]. Aspartate and glutamate are negative; lysine and arginine are positive. Histidine's imidazole pKa near 6 means it can flip between uncharged and positively charged depending on its local environment, which is one reason it appears so often in enzyme active sites [1].

### What counts as a conservative substitution?

A conservative substitution replaces one amino acid with another from the same property group, such as lysine to arginine or aspartate to glutamate. BLOSUM62 rewards these pairs with positive scores: Lys-Arg 2, Asp-Glu 2, Ile-Leu 2, Ser-Thr 1. The label is about chemistry, not about whether the change is harmless in your specific protein.

### How do I read a BLOSUM62 score?

A positive score means the pair is seen in conserved protein blocks more often than chance, and a negative score means it is seen less often [6][7]. Identity scores run from 4 to 11, and most off-diagonal pairs are negative. Use the sign and magnitude as a rough guide to how tolerant an alignment is at that position, not as a prediction of function.

### Why does Cys to Ser score negative when both are small and polar?

Cysteine and serine look similar in size and polarity, but cysteine can form disulfide bonds and serine cannot [1]. The matrix assigns Cys-Ser a score of -1, probably because a cysteine that participates in a disulfide is rarely replaced in conserved protein families. The score reflects observed substitutions, not side chain chemistry alone.

## References

1. [Berg et al. Biochemistry 8th ed., Section 2.1 Proteins are built from 20 amino acids (Macmillan digital edition)](https://digfir-published.macmillanusa.com/berg8e/berg8e_ch02_2.html)
2. [Berg, Tymoczko & Stryer. Biochemistry 5th ed., Section 3.1 (NCBI Bookshelf)](https://www.ncbi.nlm.nih.gov/books/NBK22379/)
3. [IUPAC-IUB Nomenclature and Symbolism for Amino Acids and Peptides (3AA-1, 3AA-2)](https://iupac.qmul.ac.uk/AminoAcid/AA1n2.html)
4. [Grimsley, Scholtz & Pace 2009. A summary of the measured pK values of the ionizable groups in folded proteins. Protein Sci 18:247](https://doi.org/10.1002/pro.19)
5. [Fitch et al. 2015. Arginine: its pKa value revisited. Protein Sci 24:752](https://doi.org/10.1002/pro.2647)
6. [Henikoff & Henikoff 1992. Amino acid substitution matrices from protein blocks. PNAS 89:10915](https://doi.org/10.1073/pnas.89.22.10915)
7. [NCBI BLAST matrices: BLOSUM62](https://ftp.ncbi.nih.gov/blast/matrices/BLOSUM62)
8. [Berg et al. Biochemistry 8th ed., Section 2.6 (amino acid structural preferences)](https://digfir-published.macmillanusa.com/berg8e/berg8e_ch02_7.html)
9. [Berg, Tymoczko & Stryer. Biochemistry 5th ed., Section 3.2 Primary Structure (NCBI Bookshelf)](https://www.ncbi.nlm.nih.gov/books/NBK22364/)

## Related Articles

- [Amino Acid Structure](/blog/guides/amino-acid-structure)
- [Glycine Structure](/blog/guides/glycine-structure)
- [RNA Codon Chart](/blog/guides/rna-codon-chart)
- [Disulfide Bond: Formation, Structure and Function](/knowledge/molecular-biology/disulfide-bond-formation-structure-and-function)
- [How to Calculate the Isoelectric Point (pI) of an Amino Acid or Protein](/blog/research-skills/how-to-calculate-isoelectric-point)
- [The Hydrophobic Effect: Why Proteins Fold](/blog/research-skills/hydrophobic-effect-and-protein-folding)