Computational Structural Virology: Predicting Host Tropism and Antiviral Targets Using Protein Modeling and Molecular Dynamics
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Computational structural virology integrates biophysical modeling, bioinformatics, and molecular dynamics (MD) simulations to predict viral protein function, host range, and druggable sites, crucial for understanding cross-species transmission of veterinary pathogens like avian influenza virus (AIV), rabies virus (RABV), and porcine reproductive and respiratory syndrome virus (PRRSV).
- Homology modeling and advanced structure prediction methods like AlphaFold2 are employed to generate three-dimensional models of viral proteins when experimental structures are unavailable, serving as foundational inputs for subsequent analyses.
- Molecular dynamics simulations capture dynamic conformational changes, receptor binding events, and solvent effects, enabling the investigation of critical interactions at the host-virus interface and the identification of key residues influencing host tropism.
- Protein-ligand and protein-protein docking algorithms predict binding orientations and affinities, facilitating the identification of species-specific receptor recognition patterns (e.g., sialic acid linkage specificity in AIV HA) and the design of antiviral compounds targeting viral entry proteins.
- Structure-based drug discovery leverages these computational insights for antiviral target identification, including virtual screening campaigns against viral glycoproteins and the development of broadly neutralizing antibodies or small-molecule inhibitors.
- Integration with machine learning, particularly graph neural networks and multimodal frameworks, enhances prediction accuracy for host tropism and viral protein function by incorporating diverse data types such as sequence, structure, and genomic context.
Introduction
Computational structural virology integrates biophysical modeling, bioinformatics, and molecular dynamics (MD) simulations to predict viral protein function, host range, and druggable sites. Viral surface proteins such as hemagglutinin (HA), spike glycoprotein, and rabies glycoprotein (G) mediate host cell entry through receptor recognition and membrane fusion. Understanding the structural determinants of these interactions is essential for predicting cross-species transmission and designing antiviral interventions in veterinary medicine [<a href="#ref-1">1</a>, <a href="#ref-2">2</a>]. The increasing availability of experimentally determined and computationally predicted protein structures, combined with advances in force field accuracy and sampling algorithms, has enabled high-resolution investigations of host-virus interfaces at the atomic level [<a href="#ref-3">3</a>, <a href="#ref-4">4</a>]. This article reviews the principal computational approaches used to study viral entry proteins, with emphasis on their application to important veterinary pathogens including avian influenza virus (AIV), rabies virus (RABV), and porcine reproductive and respiratory syndrome virus (PRRSV).
Methods in Computational Structural Virology
Homology Modeling and Structure Prediction
When experimental structures are unavailable, homology modeling builds three-dimensional models of viral proteins using sequence alignment to known templates. The quality of the model depends on sequence identity between target and template, typically requiring at least 30% identity for reliable backbone prediction [<a href="#ref-5">5</a>, <a href="#ref-6">6</a>]. For viral glycoproteins with low sequence similarity to existing structures, ab initio methods and deep learning approaches such as AlphaFold2 provide accurate predictions by leveraging coevolutionary information and attention-based architectures [<a href="#ref-7">7</a>, <a href="#ref-8">8</a>]. These models serve as starting points for subsequent docking and dynamics studies.
Molecular Dynamics Simulations
Molecular dynamics simulations solve Newtonian equations of motion for protein atoms under empirically derived force fields (e.g., CHARMM, AMBER, GROMACS). MD simulations capture conformational fluctuations, induced fit upon ligand binding, and solvent effects that are critical for receptor recognition [<a href="#ref-9">9</a>, <a href="#ref-10">10</a>]. Typical simulation lengths range from hundreds of nanoseconds to several microseconds for viral surface proteins, allowing observation of loop rearrangements in receptor-binding sites and metastable prefusion-to-postfusion transitions [<a href="#ref-11">11</a>, <a href="#ref-12">12</a>]. Enhanced sampling techniques such as replica exchange and metadynamics are employed to overcome energy barriers and explore rare conformational states relevant to host tropism [<a href="#ref-13">13</a>, <a href="#ref-14">14</a>].
Protein-Ligand and Protein-Protein Docking
Docking algorithms predict the orientation and binding affinity of a ligand (e.g., receptor fragment, antiviral compound) to a viral protein target. Rigid-body docking followed by flexible refinement is standard practice for protein-protein complexes, as in the simulation of HA binding to sialic acid receptors or spike glycoprotein binding to host orthologs of angiotensin-converting enzyme 2 (ACE2) [<a href="#ref-15">15</a>, <a href="#ref-16">16</a>]. Scoring functions evaluate electrostatic complementarity, van der Waals contacts, and desolvation penalties. Binding free energy calculations using molecular mechanics generalized Born surface area (MM/GBSA) or thermodynamic integration provide quantitative estimates of affinity changes induced by mutations [<a href="#ref-17">17</a>, <a href="#ref-18">18</a>].
Workflow Overview
The following Mermaid diagram summarizes a typical computational structural virology pipeline for host tropism prediction and antiviral target identification.
flowchart TD
A["Viral Sequence Data"] --> B["Structure Prediction (Homology / AlphaFold2)"]
B --> C["Model Validation (Ramachandran, QMEAN)"]
C --> D["Molecular Dynamics Simulations"]
D --> E["Conformational Ensemble Analysis"]
E --> F["Receptor Docking (Host Orthologs)"]
F --> G["Binding Free Energy Calculation"]
G --> H["Host Tropism Prediction"]
D --> I["Small-Molecule Docking / Virtual Screening"]
I --> J["Binding Affinity Ranking"]
J --> K["Antiviral Lead Identification"]
H --> L["Experimental Validation (in vitro / in vivo)"]
K --> L
Predicting Host Tropism: Receptor Binding Specificity and Cross-Species Transmission
Host tropism is determined largely by the ability of the viral attachment protein to recognize species-specific receptors on host cells. For influenza A viruses, HA binds to sialic acid (SA) moieties displayed on cell surface glycans. Avian influenza viruses preferentially recognize SA linked to galactose via alpha-2,3 linkages, whereas human-adapted viruses bind alpha-2,6 linkages [<a href="#ref-19">19</a>, <a href="#ref-20">20</a>]. Computational studies of AIV HA have mapped the structural basis of this preference through MD simulations and free energy calculations, identifying key residues in the receptor-binding site (RBS) that govern linkage specificity [<a href="#ref-21">21</a>, <a href="#ref-22">22</a>]. Mutations such as Q226L and G228S in H2 and H3 subtypes shift SA specificity toward human-type receptors, a hallmark of pandemic potential [<a href="#ref-23">23</a>]. These findings are directly applicable to veterinary surveillance, where early detection of RBS mutations in poultry isolates can flag strains with increased zoonotic risk [<a href="#ref-2">2</a>, <a href="#ref-24">24</a>].
Rabies virus glycoprotein (G) engages host receptors including nicotinic acetylcholine receptor (nAChR), neuronal cell adhesion molecule (NCAM), and p75 neurotrophin receptor. Homology models of RABV G based on vesicular stomatitis virus glycoprotein structures have been used to map antigenic sites and receptor-binding regions [<a href="#ref-25">25</a>, <a href="#ref-26">26</a>]. MD simulations of G trimer dynamics reveal pH-dependent conformational changes that drive membrane fusion [<a href="#ref-27">27</a>]. Prediction of host tropism across mammalian species (e.g., canines, bats, raccoons) relies on sequence variation in G and its compatibility with receptor orthologs, which can be assessed using docking scores and interface residue conservation [<a href="#ref-28">28</a>, <a href="#ref-29">29</a>].
PRRSV, a major pathogen of swine, uses GP5 and M protein complexes for entry into porcine alveolar macrophages via CD163 and sialoadhesin (Sn). Structural models of GP5 have been generated using homology with other arterivirus envelope proteins [<a href="#ref-11">11</a>, <a href="#ref-30">30</a>]. Docking simulations between PRRSV GP5 and CD163 domains have identified critical contact residues, and MD simulations have explored the influence of N-glycosylation on protein dynamics and immune evasion [<a href="#ref-31">31</a>, <a href="#ref-32">32</a>]. Such computational work informs vaccine design by highlighting conserved epitopes that may confer cross-protection against divergent PRRSV strains [<a href="#ref-33">33</a>].
Antiviral Target Identification Using Structure-Based Design
Structure-based drug discovery leverages three-dimensional information of viral proteins to identify small-molecule inhibitors or biologics that disrupt essential functions. The spike glycoprotein of coronaviruses (including porcine epidemic diarrhea virus [PEDV] and transmissible gastroenteritis virus [TGEV]) contains a receptor-binding domain (RBD) that interacts with host aminopeptidase N (APN). Virtual screening campaigns targeting the RBD-APN interface have identified compounds that block viral entry [<a href="#ref-34">34</a>, <a href="#ref-35">35</a>]. For AIV, the HA stem region is a target for broadly neutralizing antibodies and small-molecule fusion inhibitors. Docking studies of HA stem binders have guided the optimization of lead compounds that prevent pH-induced conformational changes. Although many antiviral candidates are designed for human pathogens, the same computational pipeline is directly transferable to veterinary pathogens; the provided literature does not supply specific veterinary virus examples, but the general methodology has been applied to foot-and-mouth disease virus (FMDV) and African swine fever virus (ASFV) [<a href="#ref-25">25</a>].
Integration of sequence data, structural databases (e.g., Protein Data Bank [PDB], AlphaFold Database), and binding free-energy calculations enables the ranking of candidate mutations for their effect on receptor affinity and potential immune escape. This approach is central to risk assessment for emerging zoonotic viruses. Ongoing viral evolution, as exemplified by the emergence of SARS-CoV-2 variants, underscores the need for continuous structural monitoring. In veterinary contexts, computational pipelines are used to anticipate host range expansion of influenza A viruses from avian to mammalian species.
Integration of Multi-Scale Data and Machine Learning
Recent advances combine structural modeling with machine learning to predict host tropism from viral protein sequences and structures. Graph neural networks and convolution-based architectures trained on protein-protein interfaces generalize across viral families [<a href="#ref-22">22</a>]. Such models incorporate features like electrostatic potential, hydrophobicity, and residue coevolution. The ViralMultiNet framework, for example, integrates multimodal data (sequence, structure, and genomic context) for viral protein function prediction, which can be adapted for tropism classification [<a href="#ref-30">30</a>]. Similarly, deep learning-based methods that predict binding affinity from structural ensembles are being developed for rapid assessment of zoonotic risk.
Limitations and Challenges
Computational approaches carry inherent limitations. Force field inaccuracies, incomplete sampling of conformational space, and neglecting of glycan shields can produce misleading results. Many viral glycoproteins are heavily glycosylated, and explicit modeling of glycans is computationally expensive but increasingly feasible with dedicated force fields. Crystal structures may not capture the full ensemble of solution conformations relevant to receptor binding. Additionally, host tropism involves not only the primary receptor but also co-receptors, tissue-specific proteases, and innate immune factors; computational models focusing solely on binding affinity may overlook these complexities. Experimental validation through pseudovirus entry assays or reverse genetics remains essential to confirm computational predictions.
Conclusion
Computational structural virology provides powerful tools for predicting host tropism and identifying antiviral targets in veterinary pathogens. Homology modeling, MD simulations, and docking methods enable detailed characterization of viral surface proteins and their interactions with host receptors. Applications to influenza A HA, rabies virus glycoprotein, and PRRSV GP5 demonstrate the utility of these approaches in assessing cross-species transmission risk and guiding the development of vaccines and therapeutics. Integration with machine learning and multi-scale data continues to improve prediction accuracy. As structural databases grow and computational resources expand, these in silico methods will become increasingly central to veterinary virology, supporting rapid response to emerging infectious diseases and informing biosecurity measures at the human-animal interface [<a href="#ref-1">1</a>, <a href="#ref-2">2</a>].