Structural Prediction of Bat Coronavirus Spike Proteins: Insights from AlphaFold2 and Molecular Dynamics
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- AlphaFold2 accurately predicts the three-dimensional structures of bat coronavirus spike proteins, including the receptor-binding domain (RBD), providing near-experimental resolution for proteins lacking experimental data.
- Molecular Dynamics (MD) simulations complement static structural predictions by revealing dynamic conformational changes, such as RBD opening and allosteric communication, essential for understanding viral entry mechanisms.
- Computational modeling of bat coronavirus RBDs interacting with ACE2 orthologs identifies key residues and polymorphisms that dictate cross-species transmission potential and zoonotic risk.
- Analysis of mutations within the RBD, particularly in the receptor-binding motif (RBM) loop, using both AlphaFold2 and MD simulations, can predict enhanced binding affinity to host receptors and potential for immune evasion.
- Integration of AlphaFold2 predictions with experimental structures (e.g., cryo-EM) allows for validation and refinement, crucial for understanding glycan shielding, identifying conserved epitopes for vaccine design, and assessing pandemic preparedness.
Introduction
Bats serve as reservoir hosts for a diverse array of coronaviruses, including alphacoronaviruses and betacoronaviruses that have demonstrated the capacity for cross-species transmission [<a href="#ref-1">1</a>, <a href="#ref-2">2</a>, <a href="#ref-3">3</a>, <a href="#ref-4">4</a>, <a href="#ref-5">5</a>]. The spike glycoprotein (S protein) is the primary determinant of host range and tissue tropism, mediating viral entry via receptor recognition and membrane fusion [<a href="#ref-6">6</a>, <a href="#ref-7">7</a>, <a href="#ref-8">8</a>]. Understanding the three-dimensional architecture of bat coronavirus spike proteins is therefore critical for assessing zoonotic spillover risk and for developing predictive models of host adaptation [<a href="#ref-9">9</a>, <a href="#ref-10">10</a>, <a href="#ref-11">11</a>]. However, experimental structure determination through cryo-electron microscopy (cryo-EM) or X-ray crystallography remains resource-intensive and is not feasible for every novel bat coronavirus isolate [<a href="#ref-8">8</a>, <a href="#ref-12">12</a>]. Computational approaches, particularly deep learning-based protein structure prediction and molecular dynamics (MD) simulations, have emerged as powerful alternatives for generating high-confidence structural models and for probing the dynamic behavior of spike proteins [<a href="#ref-13">13</a>, <a href="#ref-14">14</a>, <a href="#ref-15">15</a>, <a href="#ref-16">16</a>]. This article reviews the application of AlphaFold2 and MD simulations to bat coronavirus spike proteins, focusing on receptor-binding domain (RBD) architecture, ACE2 interaction dynamics, and implications for zoonotic risk assessment.
AlphaFold2 in Structural Prediction of Bat Coronavirus Spikes
AlphaFold2, a deep learning algorithm that predicts protein structures from amino acid sequences with near-experimental accuracy, has been extensively applied to viral glycoproteins [<a href="#ref-12">12</a>, <a href="#ref-13">13</a>, <a href="#ref-15">15</a>]. For bat coronaviruses, AlphaFold2 models have been used to generate full-length spike ectodomain structures, including the RBD, N-terminal domain (NTD), and the S1/S2 cleavage region [<a href="#ref-8">8</a>, <a href="#ref-16">16</a>]. The algorithm leverages multiple sequence alignment (MSA) features and an attention-based neural network to predict inter-residue distances and angles, producing a ranked set of structural models [<a href="#ref-13">13</a>, <a href="#ref-14">14</a>]. These models have been particularly valuable for sarbecoviruses and merbecoviruses where no experimental structure exists, such as for SHC014-CoV and BtKY72 [<a href="#ref-2">2</a>, <a href="#ref-16">16</a>]. Comparative analyses have shown that AlphaFold2 models of bat coronavirus RBDs recapitulate the core beta-sheet fold and receptor-binding motif (RBM) conformation observed in cryo-EM structures, with root-mean-square deviations (RMSD) typically below 2 Å over backbone atoms [<a href="#ref-8">8</a>, <a href="#ref-16">16</a>]. However, loop regions, especially the RBM loop that contacts ACE2, often exhibit higher conformational variability and require refinement through MD simulations [<a href="#ref-16">16</a>, <a href="#ref-17">17</a>]. AlphaFold2 also predicts the positioning of glycans attached to N-linked glycosylation sequons, which is critical for understanding glycan shielding and immune evasion [<a href="#ref-8">8</a>]. For a broader discussion of deep learning-based structural prediction of viral envelope glycoproteins, readers may refer to the article Deep Learning-Driven Structural Prediction of Viral Envelope Glycoproteins: Implications for Receptor Binding and Antigenic Drift.
Molecular Dynamics Simulations of Spike Dynamics
MD simulations provide atomic-level insights into the conformational flexibility and dynamics of spike proteins that are inaccessible from static experimental or predicted structures [<a href="#ref-10">10</a>, <a href="#ref-17">17</a>, <a href="#ref-18">18</a>]. All-atom MD simulations of bat coronavirus spike models, including those from AlphaFold2, are typically performed with explicit solvent and physiological ionic strength using force fields such as CHARMM or AMBER [<a href="#ref-14">14</a>, <a href="#ref-16">16</a>, <a href="#ref-18">18</a>]. Simulation timescales ranging from hundreds of nanoseconds to several microseconds have been employed to investigate RBD opening, the transition from closed (RBD-down) to open (RBD-up) conformations, and the allosteric communication between the RBD and the subdomain-2 (SD2) region [<a href="#ref-16">16</a>]. For example, MD simulations of SHC014-CoV spike variants carrying the F294L and A835D mutations revealed epistatic effects that increase RBD opening propensity, a prerequisite for ACE2 binding [<a href="#ref-16">16</a>]. The simulations also identified a conserved salt-bridge network involving the fusion peptide proximal region (FPPR) that modulates large-scale conformational changes [<a href="#ref-16">16</a>]. Free energy perturbation (FEP) and binding free energy calculations using methods like MM-GBSA have been applied to quantify the impact of RBD mutations on ACE2 binding affinity [<a href="#ref-17">17</a>, <a href="#ref-18">18</a>, <a href="#ref-19">19</a>]. These computational predictions have been validated against experimental binding assays for several bat coronaviruses, including the recently described heart-nosed bat alphacoronavirus that uses human CEACAM6 as an entry receptor [<a href="#ref-1">1</a>]. A detailed overview of MD simulation methodologies for viral spike glycoproteins can be found in the article Molecular Dynamics Simulations of Viral Spike Glycoproteins: Insights into Host Receptor Binding and Antibody Escape.
Receptor-Binding Domain and ACE2 Interactions
The RBD of bat coronavirus spike proteins interacts with host cell receptors, predominantly angiotensin-converting enzyme 2 (ACE2) for sarbecoviruses, although alternative receptors such as CEACAM6 and dipeptidyl peptidase 4 (DPP4) are used by some lineages [<a href="#ref-1">1</a>, <a href="#ref-6">6</a>, <a href="#ref-7">7</a>]. Structural prediction of the RBD-ACE2 complex is essential for assessing cross-species compatibility and zoonotic potential [<a href="#ref-9">9</a>, <a href="#ref-11">11</a>, <a href="#ref-20">20</a>]. AlphaFold2 has been used to model the RBDs of bat coronaviruses such as BtKY72, RsSHC014, and WIV1 in complex with ACE2 orthologs from humans, bats, and other mammals [<a href="#ref-2">2</a>, <a href="#ref-17">17</a>, <a href="#ref-21">21</a>]. These models identify key contact residues in the RBM, including those that form hydrogen bonds and hydrophobic interactions with ACE2 [<a href="#ref-2">2</a>, <a href="#ref-19">19</a>]. Comparative studies have shown that bat ACE2 polymorphisms at positions 31, 41, and 353, among others, critically affect binding affinity [<a href="#ref-7">7</a>]. For example, the bat ACE2 residue asparagine at position 31 (N31) is conserved across many bat species and is compatible with sarbecovirus RBD binding, whereas substitutions such as histidine or serine reduce affinity [<a href="#ref-7">7</a>]. MD simulations of these RBD-ACE2 complexes provide dynamic binding free energies and reveal that conformational fluctuations in the RBM loop can compensate for unfavorable substitutions [<a href="#ref-17">17</a>, <a href="#ref-19">19</a>]. The predictive capacity of these models was demonstrated in the identification of a critical residue motif YYDRxxG in the RBD that determines neutralization breadth of pan-sarbecovirus antibodies [<a href="#ref-22">22</a>]. Additional insights into receptor-binding dynamics and cross-species transmission are discussed in the article Computational Prediction of Zoonotic Spillover: Receptor-Binding Dynamics and Structural Modeling of Bat Coronavirus Spike Proteins.
Mutations and Zoonotic Spillover Potential
Spike protein mutations accumulated in bat coronaviruses through evolution in reservoir hosts and intermediate hosts can enhance human ACE2 binding and immune evasion [<a href="#ref-3">3</a>, <a href="#ref-8">8</a>, <a href="#ref-23">23</a>, <a href="#ref-24">24</a>]. Bioinformatic and computational approaches have identified several adaptive mutations in the RBD, such as N501Y, T487A, and K417N, that increase binding affinity for human ACE2 [<a href="#ref-3">3</a>, <a href="#ref-24">24</a>]. For bat coronaviruses, substitutions in the RBM that introduce positively charged residues or alter the loop flexibility are associated with increased spillover risk [<a href="#ref-10">10</a>, <a href="#ref-25">25</a>]. The furin cleavage site at the S1/S2 boundary, which is present in some sarbecoviruses but absent in many bat coronaviruses, has been investigated through sequence analysis and structural modeling [<a href="#ref-12">12</a>, <a href="#ref-25">25</a>]. MD simulations have shown that insertion of a furin motif can alter the proteolytic activation and fusion kinetics [<a href="#ref-12">12</a>]. Multi-epitope vaccine design studies have utilized predicted spike protein structures from bat coronaviruses to identify conserved epitopes that may confer cross-protection [<a href="#ref-23">23</a>, <a href="#ref-26">26</a>, <a href="#ref-27">27</a>, <a href="#ref-28">28</a>]. These in silico approaches rely on accurate structural models from AlphaFold2 to map epitope accessibility and stability [<a href="#ref-26">26</a>, <a href="#ref-27">27</a>]. The role of accessory proteins in coronavirus evolution and host adaptation has also been characterized genetically, though structural predictions for these smaller proteins remain less common [<a href="#ref-29">29</a>]. For a comprehensive treatment of spike protein evolution and receptor binding adaptation, see the article Computational Prediction of Cross-Species Receptor Binding: Bat Coronavirus Spike Protein Evolution and Human Pandemic Risk.
Comparison with Experimental Structures
Validation of predicted structures against experimentally determined cryo-EM or X-ray crystallography models is a critical step in establishing the reliability of computational predictions [<a href="#ref-8">8</a>, <a href="#ref-16">16</a>]. For bat coronavirus spikes, several cryo-EM structures have been solved for close relatives of SARS-CoV-2, including WIV1, SHC014, and civet coronaviruses [<a href="#ref-8">8</a>]. These experimental structures reveal detailed features such as glycan trees, water molecule coordination networks, and the presence of a fatty acid (linoleic acid) binding pocket that stabilizes the RBD in the down conformation [<a href="#ref-8">8</a>]. AlphaFold2 predictions generally agree well with these experimental models for ordered secondary structure elements and core beta-sheets, but differences arise in flexible loops, the SD2 region, and the NTD [<a href="#ref-8">8</a>, <a href="#ref-12">12</a>]. Additionally, cryo-EM structures have resolved biliverdin binding in the NTD, a feature that is not consistently predicted by AlphaFold2 without ligand information [<a href="#ref-8">8</a>]. MD simulations initiated from AlphaFold2 models can be used to relax the structure and improve agreement with experimental data, particularly for water-mediated hydrogen bonding networks [<a href="#ref-16">16</a>]. Comparative analysis between predicted and experimental structures has informed the design of pan-sarbecovirus therapeutics and vaccines [<a href="#ref-14">14</a>, <a href="#ref-22">22</a>]. Further discussion of integrating cryo-EM and computational modeling is available in the article Structural and Evolutionary Dynamics of Coronavirus Spike Protein: Integrating Cryo-EM, Molecular Dynamics, and Phylogenetic Surveillance.
Workflow for Structural Prediction and Dynamics Analysis
The following Mermaid diagram outlines a typical computational pipeline combining AlphaFold2 and MD simulations to investigate bat coronavirus spike protein structure and dynamics.
flowchart TD
A["Bat Coronavirus Spike Sequence"] --> B["AlphaFold2 Structure Prediction"]
B --> C{"Model Evaluation"}
C -->|"pLDDT > 70"| D["Model Refinement<br>Energy Minimization"]
C -->|"pLDDT < 50"| E["Iterative Template Selection<br>or Metagenomic Assembly"]
D --> F["System Setup<br>Solvation, Ions, Glycosylation"]
F --> G["All-Atom MD Production Run<br>100 ns - 1 μs"]
G --> H["Trajectory Analysis"]
H --> I["RMSD / RMSF / Principal Component Analysis"]
H --> J["Dynamic Cross-Correlation<br>& Community Network Analysis"]
H --> K["Binding Free Energy Calculation<br>MM-GBSA / FEP"]
I --> L["Identify Conformational States<br>RBD Open/Closed Populations"]
J --> M["Allosteric Pathways<br>e.g., SD2 to RBM"]
K --> N["Binding Affinity Predictions<br>for ACE2 Orthologs"]
L --> O["Zoonotic Risk Assessment"]
M --> O
N --> O
Summary of Key Bat Coronavirus Spike Proteins Studied Using Computational Approaches
| Bat Coronavirus | Genus | RBD Features | Predicted Structures (AlphaFold2) | Experimental Structures (Cryo-EM) | Key References |
|---|---|---|---|---|---|
| SHC014-CoV | Sarbecovirus | RBM loop with two key mutations (F294L, A835D) | Full ectodomain; MD simulations show epistasis | Partial cryo-EM (unpublished) | [<a href="#ref-3">3</a>, <a href="#ref-16">16</a>] |
| WIV1 | Sarbecovirus | High homology to SARS-CoV; linoleic acid binding pocket | RBD-ACE2 complex | Cryo-EM at 3.5 Å (2024) | [<a href="#ref-8">8</a>] |
| BtKY72 | Sarbecovirus | Unique RBD residues; bat ACE2 compatibility | RBD-ACE2 complex with multiple orthologs | None available | [<a href="#ref-2">2</a>] |
| HKU5-CoV-2 | Merbecovirus | S1 subunit targeting; DPP4 binding | RBD and S1 domain; drug docking | None available | [<a href="#ref-18">18</a>, <a href="#ref-23">23</a>] |
| MERS-CoV related (bat) | Merbecovirus | Conserved antigenic sites | Spike monomer and trimer models | Partial cryo-EM for MERS-CoV | [<a href="#ref-30">30</a>, <a href="#ref-31">31</a>] |
| Heart-nosed bat alphacoronavirus | Alphacoronavirus | Uses CEACAM6 receptor | Predicted RBD | None available | [<a href="#ref-1">1</a>] |
Conclusion
The integration of AlphaFold2-based structure prediction with all-atom molecular dynamics simulations provides a robust framework for characterizing bat coronavirus spike proteins in the absence of experimental structures [<a href="#ref-13">13</a>, <a href="#ref-14">14</a>, <a href="#ref-16">16</a>]. These computational approaches enable the identification of critical residues governing receptor specificity, the dynamic conformational changes required for viral entry, and the evolutionary pathways that may facilitate zoonotic spillover [<a href="#ref-2">2</a>, <a href="#ref-3">3</a>, <a href="#ref-10">10</a>, <a href="#ref-18">18</a>, <a href="#ref-24">24</a>]. The continued development of deep learning models and enhanced sampling techniques in MD will further improve the accuracy and predictive power of these methods [<a href="#ref-15">15</a>, <a href="#ref-28">28</a>]. Future efforts should focus on expanding coverage to understudied bat coronavirus lineages, incorporating glycan heterogeneity into simulations, and coupling structural predictions with machine learning-based risk assessment [<a href="#ref-1">1</a>, <a href="#ref-4">4</a>, <a href="#ref-23">23</a>, <a href="#ref-32">32</a>, <a href="#ref-33">33</a>, <a href="#ref-34">34</a>, <a href="#ref-35">35</a>]. Such integrated computational virology approaches will be instrumental in pandemic preparedness and the design of broad-spectrum countermeasures [<a href="#ref-22">22</a>, <a href="#ref-26">26</a>, <a href="#ref-27">27</a>, <a href="#ref-30">30</a>].