Levinthal's Paradox and Anfinsen's Experiment: How Proteins Find Their Fold
By Dr. Zubair Khalid, DVM, MS, PhD ·

A polypeptide chain emerging from the ribosome has to become a specific three-dimensional object, and it usually does so in seconds or less. Levinthal's paradox is the observation that a random search through all possible conformations would take longer than the age of the universe, so folding cannot be random. Anfinsen's experiment on ribonuclease supplied the other half of the story: the amino acid sequence itself contains enough information to specify the native structure, with no template and no external instructions required.
You will meet this material in biochemistry coursework, in structural biology seminars, and in any discussion of misfolding disease or protein engineering. It also comes up whenever someone claims that a folding algorithm "solves" the problem, or that a molecular dynamics trajectory "proves" a folding pathway. Knowing what Anfinsen actually showed, and what Levinthal's estimate actually assumes, keeps those conversations honest.
Quick Answer
- Levinthal's paradox: if a chain found its native state by randomly sampling every possible conformation, the search would take an absurdly long time, yet real proteins fold in seconds or less [4].
- Anfinsen's thermodynamic hypothesis: the native structure is the conformation in which the Gibbs free energy of the whole system is lowest, and that conformation is determined by the amino acid sequence in a given environment [2].
- The classic estimate: a 100-residue chain with 3 conformations per residue has about $3^{100} \approx 5 \times 10^{47}$ configurations; at $10^{13}$ conformations per second that is roughly $10^{27}$ years [4].
- The resolution: folding is not a random search. A small energetic bias against locally incorrect conformations, on the order of a few $kT$, collapses the search time to about a second [4].
- The experimental anchor: reduced, unfolded ribonuclease refolds spontaneously when denaturant is removed, and the product is virtually identical to the native enzyme [8].
- The practical reading: sequence determines structure, but the route to that structure is guided by the energy surface, not by exhaustive sampling.
Anfinsen's Experiment and the Thermodynamic Hypothesis
Bovine pancreatic ribonuclease is a small, well-behaved test case. It contains 124 amino acid residues and four disulfide bonds [2]. Anfinsen's group, working with Michael Sela and Fred White in 1956 and 1957, treated the enzyme with beta-mercaptoethanol in 8 M urea. That combination breaks the disulfides and unfolds the chain, giving a fully reduced, randomly coiled polypeptide with no enzymatic activity [8].
The interesting part came next. When urea and beta-mercaptoethanol were removed by dialysis, the reduced enzyme slowly regained activity as its sulfhydryl groups oxidized in air. The refolded protein was virtually identical to native ribonuclease [8]. No cellular machinery, no template, no chaperone: just the sequence, the solvent, and time.
The disulfide arithmetic makes the result sharper. Eight cysteines can be paired into four disulfides in
$$(2n-1)(2n-3)\cdots 1 = 7 \times 5 \times 3 \times 1 = 105$$
ways, where $n = 4$ is the number of disulfide bonds. Only one of those 105 combinations is enzymatically active [8][2]. If pairing were random, you would expect the native arrangement about 1 time in 105, or 0.95 percent. Anfinsen's scrambled ribonuclease, produced by reoxidizing the reduced protein while it was still in 8 M urea, is a mixture of many or all of the 105 isomers with on the order of 1 percent of native activity [2]. The match between 1/105 and the observed 1 percent is a coincidence of a simple model, not proof that pairing is random.
The decisive control was the third experiment. When urea was removed and scrambled ribonuclease was exposed to a small amount of mercaptoethanol, disulfide interchange converted the mixture into a homogeneous product indistinguishable from native ribonuclease. The driving force was entirely the free energy gained in reaching the native structure [2]. That is the thermodynamic hypothesis in one experiment: the native conformation is the one in which the Gibbs free energy of the whole system is lowest, and it is determined by the totality of interatomic interactions, hence by the amino acid sequence, in a given environment [2].
Anfinsen was careful about the environment. A protein only makes stable structural sense under conditions similar to those for which it was selected, the physiological state [2]. Change the pH, the ionic strength, the metal ions or the temperature far enough and the same sequence can adopt a different, often nonfunctional, conformation.
Anfinsen received half of the 1972 Nobel Prize in Chemistry for this work on ribonuclease and the connection between amino acid sequence and biologically active conformation [3]. Stanford Moore and William H. Stein shared the other half for work on the connection between chemical structure and catalytic activity of the active center of the ribonuclease molecule [3]. Anfinsen delivered his Nobel Lecture, "Studies on the principles that govern the folding of protein chains," on December 11, 1972 [2], and the printed version appeared in Science in 1973 [1].
Levinthal's Paradox: The Arithmetic of a Random Search
The paradox is usually stated as a counting problem. Take a chain of 101 residues, so there are 100 bonds connecting them. Give each bond three possible states. The number of configurations is then
$$N_{\text{conf}} = 3^{100} \approx 5 \times 10^{47}$$
where $N_{\text{conf}}$ is the total number of conformations and the exponent counts the bonds. Sample conformations at $10^{13}$ per second, which is roughly $3 \times 10^{20}$ per year, and trying them all takes about $10^{27}$ years [4]. Berg states the same estimate as $5 \times 10^{34}$ seconds, or $1.6 \times 10^{27}$ years, for a 100-residue protein with three conformations per residue, while noting that small proteins can fold in less than a second [8].
The number is deliberately crude. Three states per residue and $10^{13}$ samples per second are assumptions, and the answer moves by orders of magnitude when you change them. Two states per residue gives about $4.0 \times 10^{9}$ years. Three states for 150 residues gives about $1.2 \times 10^{51}$ years. The conclusion does not change, which is the point of the illustration.
One historical detail matters for citation hygiene. Levinthal's estimate appears in his 1969 contribution to the proceedings Mössbauer Spectroscopy in Biological Systems (University of Illinois Press, pages 22 to 24). The frequently cited 1968 J Chim Phys paper, "Are there pathways for protein folding?", does not contain the paradox, according to Zwanzig, Szabo and Bagchi [4][5]. Cite the 1968 paper for the pathways question, and attribute the numbers through Zwanzig et al. or Berg.
Worked Example
This is a thought experiment, not a model of real folding. Assume a 100-residue chain (Zwanzig et al. phrase it as 101 residues with 100 bonds), 3 conformations per residue, and $10^{13}$ conformations sampled per second.
N = 3.0**100 # 5.154e47 conformations
t = N / 1e13 # 5.154e34 seconds
years = t / (365.25*24*3600) # 1.633e27 years
The result is $5.154 \times 10^{47}$ conformations, $5.154 \times 10^{34}$ seconds, and $1.633 \times 10^{27}$ years. That matches Berg's $5 \times 10^{34}$ seconds and $1.6 \times 10^{27}$ years, and Zwanzig et al.'s "about $10^{27}$ years" [4][8].
Sensitivity is easy to check. Two states per residue gives $1.268 \times 10^{30}$ conformations and $4.0 \times 10^{9}$ years. Ten states gives $1 \times 10^{100}$ and $3.2 \times 10^{79}$ years. Three states with 150 residues gives $3.7 \times 10^{71}$ and $1.2 \times 10^{51}$ years.
Now add a bias. Zwanzig, Szabo and Bagchi used the approximation
$$T = \frac{1}{N k_0}\left(1 + \frac{k_0}{k_1}\right)^N$$
where $T$ is the mean first-passage time to the native state, $N$ is the number of bonds, $k_0$ is the rate of forming an incorrect local configuration, and $k_1$ is the rate of forming a correct one. With $N = 100$, $k_1 = 10^{9}$ per second, and $k_0 = 2\exp(-U/kT) \times 10^{9}$ per second, where $U$ is the energy penalty for an incorrect bond and $kT$ is the thermal energy, the numbers run as follows. At $U = 0$, $T = 2.6 \times 10^{36}$ seconds. At $U = 1\,kT$, $T = 1.2 \times 10^{13}$ seconds, about 380,000 years. At $U = 2\,kT$, $T = 0.94$ seconds. At $U = 3\,kT$, $T = 1.3 \times 10^{-6}$ seconds. A bias of only about $2\,kT$ per bond turns an astronomically long search into about a second, which is the paper's point [4].
Two cautions on the formula. The worked example uses the approximate Eq. 1, which the authors say gives slightly smaller times at large $U/kT$; it even falls below their fully biased limit of about $5 \times 10^{-9}$ seconds once $U$ reaches about $5\,kT$ (the approximate formula gives about $1.0 \times 10^{-8}$ seconds at $4\,kT$ and $2.8 \times 10^{-9}$ seconds at $5\,kT$). Quote only up to about $3\,kT$. Without any bias, the same expression reduces to
$$T = \frac{1}{N k_0}(v+1)^N$$
the formula usually used in Levinthal discussions, where $(v+1)^N$ is the number of configurations [4].
For the disulfide side of Anfinsen's experiment, the same counting logic applies. Eight cysteines pair in 105 ways, so random reoxidation would give the native pairing about 1/105, or 0.95 percent of the time, in line with the roughly 1 percent activity of scrambled ribonuclease reported by Anfinsen [2].
How Proteins Actually Find the Fold
The resolution is cumulative selection: folding proceeds by retaining partly correct intermediates instead of sampling all conformations at random [8]. Each favorable local contact narrows the set of conformations still worth exploring, so the search is biased toward the native state from the start.
The energy landscape picture makes this concrete. Picture a funnel whose wide rim represents the many structures accessible to the ensemble of denatured molecules, and whose narrow base is the native state [8]. Dill and Chan described the "new view" of folding as parallel, multi-pathway, diffusion-like processes in which folding funnels to a single stable state by multiple routes, which removes the paradox [6]. They also pointed out that the classical pathway idea, used to solve the needle-in-a-haystack search, seemed to conflict with Anfinsen's finding that folding is pathway-independent [6]. If the native state is the free energy minimum, many routes can reach it, and no single obligatory sequence of intermediates is required.
Rates support the funnel picture. Proteins without disulfide cross-links, such as staphylococcal nuclease or myoglobin, can renature almost completely in a few seconds or less [2]. Staphylococcal nuclease refolds in at least two phases: fast nucleation and folding with a half-time of about 50 milliseconds, then a slower step with a half-time of about 200 milliseconds [2].
Disulfide-containing proteins complicate the timing. In vitro renaturation of reduced ribonuclease was slow, often hours, while synthesis of the 124-residue chain in tissues takes about 2 minutes [2]. That discrepancy led to the discovery of an endoplasmic reticulum enzyme that catalyzes disulfide interchange and forms the native pairing in less than the requisite two minutes [2].
Inside cells, the situation is harder still. Newly synthesized proteins are at great risk of misfolding and aggregation, and cells invest in a network of molecular chaperones that prevent aggregation and promote efficient folding [7]. Chaperones block undesirable interactions, such as aggregation, that would otherwise compete with correct folding [8][7]. Because proteins are highly dynamic, constant chaperone surveillance is needed to maintain protein homeostasis, or proteostasis, and an age-related decline in proteostasis capacity is linked to aggregation diseases such as Alzheimer's and Parkinson's disease [7].
Common Mistakes
- Treating $10^{27}$ years as a measured quantity. It is an illustrative estimate built on assumed values for states per residue and sampling rate. Change the assumptions and the number moves by orders of magnitude, though the conclusion holds.
- Confusing Anfinsen's dogma with a universal law. "Anfinsen's dogma" is a later popular label; Anfinsen himself called it the thermodynamic hypothesis. Commonly cited exceptions, not reviewed here, include metastable proteins such as serpins, prions, and proteins whose folding requires chaperones or pro-sequences.
- Reading the 1 percent activity of scrambled ribonuclease as proof of random pairing. The 1/105 calculation is a simple model, and the agreement with the observed value is a coincidence, not evidence about the mechanism.
- Assuming the native state is always the global free energy minimum under any conditions. The thermodynamic hypothesis is stated for the normal physiological milieu: solvent, pH, ionic strength, metal ions or prosthetic groups, and temperature [2].
- **Citing the 1968 J Chim Phys paper for the paradox.** The estimate is in Levinthal's 1969 Allerton House proceedings chapter; the 1968 paper is about pathways [4][5].
- Forgetting that in vitro refolding can be slow. Reduced ribonuclease takes hours to refold in a test tube, while the same chain is synthesized in about 2 minutes in tissue, a discrepancy that led to the discovery of an enzyme that catalyzes disulfide interchange [2].
Limitations
The Levinthal estimate is a deliberately crude illustration. The 3 states per residue and $10^{13}$ per second rates are assumptions, and the resulting time depends strongly on them. Label all such numbers as illustrative.
The Zwanzig, Szabo and Bagchi result has its own boundaries. Their "about 1 second at $2\,kT$" figure comes from their exact Eq. 16; the worked example here uses the approximate Eq. 1, which the authors say gives slightly smaller times at large $U/kT$ and can fall below their fully biased limit. Quote only up to about $3\,kT$ [4].
The historical record has gaps. Levinthal's original numbers appeared in a 1969 conference proceedings chapter that is hard to obtain, so they are usually cited through Zwanzig et al. or Berg [4][8]. Anfinsen's own account is most accessible in his Nobel lecture on nobelprize.org [1][2].
The thermodynamic hypothesis has documented exceptions that were not researched here. Metastable proteins, prions, and proteins that require chaperones or pro-sequences to fold correctly all complicate the simple statement that sequence plus environment equals native structure. Treat the hypothesis as the foundation of the field, not as a claim with no exceptions.
Frequently Asked Questions
What is the Levinthal paradox in simple terms?
If a protein had to find its native shape by trying every possible conformation at random, the search would take longer than the age of the universe. Real proteins fold in seconds or less, so the search cannot be random [4]. The paradox is resolved by recognizing that folding is biased toward the native state by local energetics, not by exhaustive sampling.
What did the Anfinsen experiment actually show?
Anfinsen and colleagues showed that reduced, unfolded ribonuclease refolds spontaneously when urea and beta-mercaptoethanol are removed, regaining activity and a structure virtually identical to the native enzyme [8]. The experiment demonstrated that the amino acid sequence contains the information needed to specify the native conformation, which is the thermodynamic hypothesis [2].
How does the protein folding funnel resolve the paradox?
The funnel picture replaces a flat, random search with a tilted energy surface. The wide rim represents the many conformations available to the denatured ensemble, and the narrow base is the native state [8]. Because the surface slopes downward toward the native state, folding proceeds by multiple parallel routes and does not require sampling every possibility [6].
Why do proteins fold quickly in cells but slowly in a test tube?
Proteins without disulfide bonds can refold in seconds or less in vitro [2]. Disulfide-containing proteins such as ribonuclease can take hours in a test tube, while the same chain is synthesized in about 2 minutes in tissue, a discrepancy that led to the discovery of an endoplasmic reticulum enzyme that catalyzes disulfide interchange [2]. Inside cells, chaperones also prevent aggregation and promote efficient folding [7].
Do chaperones change the folding code?
No. Chaperones do not encode structural information; they block undesirable interactions such as aggregation that would otherwise compete with correct folding [8][7]. The native structure is still determined by the amino acid sequence and the environment, in line with the thermodynamic hypothesis [2].
References
- Anfinsen 1973. Principles that govern the folding of protein chains. Science 181:223
- Anfinsen, Nobel Lecture 1972: Studies on the principles that govern the folding of protein chains (PDF)
- The Nobel Prize in Chemistry 1972 (summary)
- Zwanzig, Szabo & Bagchi 1992. Levinthal's paradox. PNAS 89:20
- Levinthal 1968. Are there pathways for protein folding? J Chim Phys 65:44
- Dill & Chan 1997. From Levinthal to pathways to funnels. Nat Struct Biol 4:10
- Hartl, Bracher & Hayer-Hartl 2011. Molecular chaperones in protein folding and proteostasis. Nature 475:324
- Berg et al. Biochemistry 8th ed., Section 2.6 The amino acid sequence of a protein determines its 3D structure
- Anfinsen, Haber, Sela & White 1961. Kinetics of formation of native ribonuclease during oxidation of the reduced chain. PNAS 47:1309