Protein Structure: Biophysical Levels of Folding, Force Fields, and Conformational Stability
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Protein structure is hierarchically organized from primary (amino acid sequence) to quaternary (subunit assembly), with secondary (alpha-helices, beta-sheets) and tertiary (3D fold) structures arising from physical forces like hydrogen bonding, electrostatic interactions, and the hydrophobic effect.
- Force fields, including classical (AMBER, CHARMM), coarse-grained (MARTINI), and quantum mechanical/machine learning potentials, are essential for molecular modeling and simulating protein dynamics, folding pathways, and conformational stability.
- Conformational stability, quantified by Gibbs free energy of folding (ΔG_fold), is determined by the thermodynamic balance of enthalpy and entropy, with the hydrophobic effect being a major driving force; experimental methods like CD and DSC, alongside computational predictions (e.g., ΔΔG), characterize this stability.
- Modern deep learning approaches, exemplified by AlphaFold2 and ESMFold, have revolutionized tertiary structure prediction, achieving near-atomic accuracy and enabling the rapid generation of vast protein structure databases.
- Understanding protein folding and stability is critical for veterinary medicine, informing vaccine design against viral glycoproteins, predicting the impact of mutations on viral evolution and immune evasion, and developing structure-based diagnostics and therapeutics.
- Beyond stable globular structures, intrinsically disordered proteins (IDPs) and phase-separated states represent important functional conformations, often regulated by multivalent interactions and capable of folding-upon-binding.
Introduction
Proteins are the molecular machines of life, and their three-dimensional structures dictate their biological functions. The process by which a linear polypeptide chain adopts a unique, functional conformation is known as protein folding. This process is governed by a complex interplay of physical forces, including van der Waals interactions, hydrogen bonding, electrostatic forces, and the hydrophobic effect [<a href="#ref-1">1</a>, <a href="#ref-2">2</a>]. Understanding the biophysical basis of protein folding is essential for interpreting structure-function relationships, predicting the effects of mutations, and designing therapeutic interventions in veterinary medicine and diagnostics [<a href="#ref-3">3</a>, <a href="#ref-4">4</a>]. This article provides an exhaustive review of the hierarchical levels of protein folding, the force fields used to model them, and the thermodynamic and kinetic principles that determine conformational stability.
Hierarchical Levels of Protein Folding
Protein structure is traditionally described at four levels: primary, secondary, tertiary, and quaternary. However, modern biophysics also recognizes the importance of intrinsically disordered regions and phase-separated states [<a href="#ref-5">5</a>, <a href="#ref-6">6</a>].
Primary Structure
The primary structure is the linear sequence of amino acids linked by peptide bonds. This sequence encodes all the information necessary for folding [<a href="#ref-7">7</a>, <a href="#ref-8">8</a>]. Mutations in the primary sequence can alter folding pathways and stability, as demonstrated in studies of viral envelope proteins where single amino acid substitutions affect receptor binding and immune evasion [<a href="#ref-9">9</a>, <a href="#ref-10">10</a>].
Secondary Structure
Local hydrogen bonding patterns give rise to regular secondary structures: alpha-helices, beta-sheets, and turns. These elements are stabilized by backbone amide-carbonyl interactions [<a href="#ref-2">2</a>, <a href="#ref-11">11</a>]. Secondary structure prediction algorithms, such as PSIPRED, use neural networks to assign these elements from sequence alone [<a href="#ref-12">12</a>]. The accuracy of such predictions has improved dramatically with deep learning [<a href="#ref-13">13</a>].
Tertiary Structure
Tertiary structure refers to the global three-dimensional arrangement of a single polypeptide chain. It is stabilized by side-chain interactions, including hydrophobic packing, salt bridges, disulfide bonds, and hydrogen bonds [<a href="#ref-1">1</a>, <a href="#ref-14">14</a>]. The folding process involves a cooperative transition from an unfolded ensemble to a compact native state, often described by a funnel-shaped energy landscape [<a href="#ref-4">4</a>, <a href="#ref-15">15</a>]. Experimental techniques such as cryo-electron microscopy (cryo-EM) and X-ray crystallography provide atomic-resolution structures [<a href="#ref-16">16</a>, <a href="#ref-17">17</a>]. Computational methods like homology modeling (e.g., MODELLER, SWISS-MODEL) and deep learning (e.g., AlphaFold) have revolutionized tertiary structure prediction [<a href="#ref-7">7</a>, <a href="#ref-18">18</a>, <a href="#ref-19">19</a>].
Quaternary Structure
Many proteins function as oligomeric complexes. Quaternary structure describes the arrangement of multiple subunits. The assembly is driven by complementary surfaces and often involves allosteric regulation [<a href="#ref-14">14</a>, <a href="#ref-20">20</a>]. For example, the RACK1 scaffolding protein adopts a seven-bladed beta-propeller that mediates protein-protein interactions in signaling pathways [<a href="#ref-20">20</a>]. Bacterial microcompartment shell proteins self-assemble into polyhedral shells through charged residue interactions [<a href="#ref-14">14</a>].
Intrinsically Disordered Regions and Phase Separation
Not all proteins fold into stable globular structures. Intrinsically disordered proteins (IDPs) lack persistent secondary or tertiary structure under physiological conditions [<a href="#ref-5">5</a>, <a href="#ref-6">6</a>]. They often undergo folding-upon-binding when interacting with partners [<a href="#ref-6">6</a>]. Liquid-liquid phase separation, driven by multivalent interactions and RNA G-quadruplexes, is a key mechanism for organizing cellular compartments [<a href="#ref-5">5</a>]. The conformational dynamics of IDPs can be modulated by small molecules, as shown for the SARS-CoV-2 nucleocapsid protein [<a href="#ref-21">21</a>].
Force Fields in Molecular Modeling
Molecular dynamics (MD) simulations rely on force fields to compute the potential energy of a protein system as a function of atomic coordinates [<a href="#ref-11">11</a>, <a href="#ref-22">22</a>]. Force fields are parameterized sets of equations that describe bonded interactions (bonds, angles, dihedrals) and non-bonded interactions (van der Waals and electrostatic) [<a href="#ref-2">2</a>, <a href="#ref-23">23</a>].
Classical Force Fields
Classical force fields such as AMBER, CHARMM, and OPLS use fixed partial charges and harmonic potentials. They are widely used for simulating protein dynamics and folding [<a href="#ref-11">11</a>, <a href="#ref-22">22</a>]. The accuracy of these force fields depends on the parameterization of dihedral angle potentials and the treatment of solvation [<a href="#ref-2">2</a>]. Recent developments include polarizable force fields that account for electronic polarization [<a href="#ref-23">23</a>].
Coarse-Grained Force Fields
For large systems or long timescales, coarse-grained models reduce computational cost by grouping atoms into beads. The MARTINI force field is a popular example. Coarse-grained simulations have been applied to study protein self-assembly and membrane interactions [<a href="#ref-14">14</a>, <a href="#ref-24">24</a>].
Quantum Mechanical and Hybrid Methods
Quantum mechanical (QM) methods provide accurate descriptions of bond breaking and formation, but are computationally expensive. Hybrid QM/MM approaches combine a QM region for the active site with a MM region for the rest of the protein [<a href="#ref-1">1</a>, <a href="#ref-22">22</a>]. These methods are essential for studying enzyme catalysis and metal-binding proteins [<a href="#ref-1">1</a>].
Machine Learning Potentials
Recent advances have introduced machine learning potentials trained on quantum chemical data. These potentials achieve near-QM accuracy at MM cost [<a href="#ref-13">13</a>, <a href="#ref-23">23</a>]. They are particularly useful for modeling conformational transitions and ligand binding [<a href="#ref-9">9</a>].
Conformational Stability
Conformational stability is the thermodynamic preference of a protein for its native state over unfolded or misfolded states. It is quantified by the Gibbs free energy of folding (ΔG_fold) [<a href="#ref-4">4</a>, <a href="#ref-15">15</a>].
Thermodynamic Principles
The folding equilibrium is governed by the balance between enthalpy (favorable interactions) and entropy (chain conformational entropy and solvent entropy). The hydrophobic effect, driven by the release of water molecules from nonpolar surfaces, is a major contributor to folding stability [<a href="#ref-2">2</a>, <a href="#ref-4">4</a>]. The stability of a protein can be perturbed by temperature, pH, and denaturants [<a href="#ref-2">2</a>].
Kinetic Aspects
Folding kinetics describe the rate at which a protein reaches its native state. Many proteins fold via a two-state mechanism, while larger proteins may populate intermediate states [<a href="#ref-6">6</a>, <a href="#ref-15">15</a>]. Chaperones assist folding by preventing aggregation [<a href="#ref-2">2</a>]. The folding landscape can be rugged, with local minima corresponding to misfolded states [<a href="#ref-4">4</a>].
Experimental Characterization
Biophysical techniques such as circular dichroism (CD), fluorescence spectroscopy, and differential scanning calorimetry (DSC) are used to monitor folding transitions [<a href="#ref-1">1</a>, <a href="#ref-2">2</a>]. For example, the unfolding of bovine serum albumin by Nb₂C nanochaperones was studied using spectroscopic methods [<a href="#ref-2">2</a>]. Surface plasmon resonance (SPR) and isothermal titration calorimetry (ITC) measure binding affinities and conformational changes [<a href="#ref-21">21</a>].
Computational Prediction of Stability
Computational methods can predict the effect of mutations on protein stability. Tools like FoldX, Rosetta, and machine learning models use energy functions to estimate ΔΔG [<a href="#ref-3">3</a>, <a href="#ref-4">4</a>]. The biophysical fitness landscape approach designs sequences that trap viral evolution by destabilizing essential proteins [<a href="#ref-4">4</a>]. Structural feature-based machine learning has been applied to predict protein-protein interface stability [<a href="#ref-3">3</a>].
Computational Structure Prediction
The field of protein structure prediction has undergone a revolution with the advent of deep learning. AlphaFold2 achieved atomic accuracy in the CASP14 experiment, solving the classical protein folding problem for single domains [<a href="#ref-7">7</a>, <a href="#ref-25">25</a>]. The AlphaFold Protein Structure Database now contains over 214 million predicted structures, covering most known protein sequences [<a href="#ref-26">26</a>, <a href="#ref-27">27</a>]. Language models such as ESMFold enable rapid structure prediction without multiple sequence alignments, achieving up to 60x speedup [<a href="#ref-8">8</a>, <a href="#ref-10">10</a>]. Other methods like I-TASSER and Phyre use threading and ab initio modeling [<a href="#ref-28">28</a>, <a href="#ref-29">29</a>]. Structure comparison tools like DALI, TM-align, and Foldseek allow searching of structural databases [<a href="#ref-30">30</a>, <a href="#ref-31">31</a>, <a href="#ref-32">32</a>]. De novo design methods such as RFdiffusion generate novel protein backbones for therapeutic applications [<a href="#ref-33">33</a>].
The following Mermaid diagram summarizes the hierarchical folding levels and the computational methods used at each stage.
flowchart TD
A["Primary Sequence"] --> B["Secondary Structure Prediction"]
B --> C["Tertiary Structure Prediction"]
C --> D["Quaternary Structure Assembly"]
D --> E["Functional Dynamics"]
B --> F["PSIPRED, Deep Learning"]
C --> G["AlphaFold, I-TASSER, MODELLER"]
D --> H["Rosetta, RFdiffusion"]
E --> I["MD Simulations, Normal Mode Analysis"]
F --> J["Force Fields: Classical, QM/MM, ML"]
G --> J
H --> J
I --> J
J --> K["Conformational Stability Analysis"]
K --> L["Experimental Validation: CD, DSC, Cryo-EM"]
Applications in Veterinary Medicine and Diagnostics
Understanding protein structure and stability is critical for veterinary virology and diagnostics. For example, predicting the structure of viral glycoproteins informs vaccine design and antibody neutralization [<a href="#ref-9">9</a>, <a href="#ref-33">33</a>]. The conformational dynamics of viral envelope proteins affect host range and immune evasion [<a href="#ref-9">9</a>]. Computational modeling of host-pathogen protein-protein interactions can identify targets for antiviral therapy [<a href="#ref-3">3</a>, <a href="#ref-4">4</a>]. In diagnostics, structure-based design of capture antibodies and antigens improves assay sensitivity [<a href="#ref-21">21</a>]. The ability to predict the impact of mutations on protein stability is essential for monitoring emerging viral variants in animal populations [<a href="#ref-4">4</a>, <a href="#ref-13">13</a>].
Conclusion
Protein folding is a hierarchical process governed by fundamental physical forces. Modern computational methods, from classical force fields to deep learning, have enabled accurate prediction of protein structures and stability. These tools are transforming veterinary structural biology, allowing researchers to understand disease mechanisms and develop novel interventions. Continued integration of biophysical principles with machine learning will further advance the field.