AI-Driven Protein Language Models for Predicting Viral Host Tropism and Zoonotic Potential
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Protein Language Models (PLMs), leveraging transformer architectures, analyze viral glycoprotein sequences to predict host tropism and zoonotic potential by learning biophysical and evolutionary constraints from large sequence corpora. These models generate dense embeddings that encode critical structural and functional information, such as receptor-binding motifs and glycosylation sites, without requiring explicit labels.
- PLMs are trained on curated host-pathogen interaction datasets, using viral sequences with known host associations (e.g., influenza HA from avian/swine/human hosts, coronavirus spike from bat/camel/mammalian hosts) as input for supervised classifiers to predict host categories or zoonotic risk scores.
- Computational validation of PLM predictions is enhanced by molecular dynamics (MD) simulations and molecular docking, which quantify binding free energies between viral surface proteins and host receptors, thereby identifying key interface residues and stabilizing complexes.
- PLM-based classifiers demonstrate superior performance and generalizability compared to traditional phylogenetic methods and classical machine learning approaches that rely on hand-crafted features, as PLMs automatically learn relevant features directly from raw sequences.
- Attention weights derived from PLMs can be mapped onto 3D protein structures, allowing for visualization of residues critical for host tropism prediction, thereby bridging sequence-based insights with structural biology for hypothesis generation regarding cross-species transmission mechanisms.
- Limitations of PLMs include potential biases in training data for less-characterized pathogens, inherent lack of interpretability beyond attention weights, and the inability to directly account for post-translational modifications or glycan shielding, necessitating integration with other data modalities and physics-based simulations for comprehensive risk assessment.
Introduction
Predicting the host range of emerging viruses is a central challenge in veterinary virology and pandemic preparedness [<a href="#ref-1">1</a>]. Zoonotic pathogens, such as influenza A viruses, coronaviruses, and hantaviruses, repeatedly cross species barriers, causing disease in domestic animals and wildlife [<a href="#ref-2">2</a>, <a href="#ref-3">3</a>]. Traditional approaches to assess host tropism rely on phylogenetic analysis, receptor-binding assays, and experimental infections [<a href="#ref-4">4</a>, <a href="#ref-5">5</a>]. However, these methods are time-consuming and may not capture the subtle molecular determinants that govern cross-species transmission. In recent years, transformer-based protein language models (PLMs) have emerged as a powerful computational framework for predicting viral host range directly from glycoprotein sequences [<a href="#ref-6">6</a>]. These models learn biophysical and evolutionary constraints from large protein sequence corpora and generate dense embeddings that encode structural and functional information [<a href="#ref-7">7</a>]. This article reviews the application of PLMs to predict zoonotic potential, with a focus on spike and surface glycoproteins, and compares their performance with conventional machine learning and phylogenetic methods.
Background on Viral Host Tropism and Zoonotic Potential
Viral host tropism is primarily determined by the interaction between viral surface proteins and specific host cell receptors [<a href="#ref-8">8</a>]. For influenza A viruses, hemagglutinin (HA) binds sialic acid receptors, and the binding specificity (α2,3 vs. α2,6 linkages) is a major determinant of avian versus mammalian tropism [<a href="#ref-2">2</a>]. For coronaviruses, the spike glycoprotein engages host receptors such as angiotensin-converting enzyme 2 (ACE2) or aminopeptidase N [<a href="#ref-4">4</a>]. Hantaviruses rely on integrin receptors for entry, and variation in the glycoprotein sequence influences host specificity [<a href="#ref-3">3</a>, <a href="#ref-5">5</a>]. Zoonotic spillover occurs when a virus acquires mutations that enable efficient binding to a new host receptor [<a href="#ref-1">1</a>, <a href="#ref-9">9</a>]. Surveillance of viral diversity in animal reservoirs, including rodents, bats, birds, and livestock, is therefore critical for risk assessment [<a href="#ref-10">10</a>, <a href="#ref-11">11</a>, <a href="#ref-12">12</a>, <a href="#ref-13">13</a>, <a href="#ref-14">14</a>]. Recent metavirome analyses have cataloged numerous viral sequences in dairy cattle, urban rats [<a href="#ref-5">5</a>], and other species, highlighting the vast potential for cross-species transmission.
Protein Language Models: Architecture and Embedding Generation
Protein language models are deep neural networks based on the transformer architecture, originally developed for natural language processing [<a href="#ref-15">15</a>]. They are trained on millions of protein sequences using a masked language modeling objective, where random amino acids are masked and the model learns to predict them from context [<a href="#ref-6">6</a>]. This training captures coevolutionary patterns, structural propensities, and functional constraints without requiring explicit labels [<a href="#ref-16">16</a>]. For viral glycoproteins, the sequence is tokenized into individual residues, passed through multiple self-attention layers, and produces per-residue embeddings [<a href="#ref-17">17</a>]. These embeddings can be aggregated (e.g., via mean pooling) to obtain a fixed-length sequence representation, or they can be used as input to downstream classifiers that exploit the full spatial information [<a href="#ref-18">18</a>]. Models such as ESM-1b and ProtBERT have been widely applied to predict variant effects, protein stability, and interactions [<a href="#ref-7">7</a>, <a href="#ref-19">19</a>]. In the context of host tropism, these embeddings encode information about receptor-binding motifs, glycosylation sites, and conserved domains that are critical for cross-species recognition [<a href="#ref-20">20</a>, <a href="#ref-21">21</a>].
Training on Curated Host-Pathogen Interaction Datasets
To predict viral host range, PLM-derived embeddings are used as features for supervised classifiers. The training data must consist of viral sequences with known host associations, such as influenza HA subtypes isolated from avian, swine, or human hosts [<a href="#ref-2">2</a>, <a href="#ref-22">22</a>]. Similarly, coronavirus spike sequences from bat, camel, and other mammalian hosts provide examples of different tropism profiles [<a href="#ref-4">4</a>, <a href="#ref-14">14</a>]. The classifier can be a simple logistic regression, a random forest, or a deep neural network that takes the embedding as input and outputs a probability distribution over host categories [<a href="#ref-23">23</a>]. Attention mechanisms allow the model to highlight which residues are most influential for the prediction, often mapping to known receptor-binding domains [<a href="#ref-24">24</a>, <a href="#ref-25">25</a>]. Cross-validation on independent datasets, such as newly discovered viruses from metavirome surveys or experimental host range data, is essential to avoid overfitting. The resulting models can assign a zoonotic risk score to a novel viral sequence, flagging those with high potential for spillover [<a href="#ref-26">26</a>]. The workflow is summarized in the following diagram.
graph TD
A["Viral Glycoprotein Sequence"] --> B["Tokenization & Embedding via PLM"]
B --> C["Sequence Embedding Vector"]
C --> D["Supervised Classifier"]
D --> E["Predicted Host Tropism"]
C --> F["Attention Weights for Key Residues"]
F --> G["Mapping to 3D Structure"]
G --> H["Visualization in 3D Protein Viewer"]
D --> I["Zoonotic Risk Score"]
I --> J["Validation with MD/Docking"]
Validation with Receptor-Binding Dynamics Simulations
A prediction of host tropism is strengthened by computational validation using molecular dynamics (MD) simulations and molecular docking [<a href="#ref-4">4</a>, <a href="#ref-27">27</a>]. For a candidate viral variant, the binding free energy between the surface protein and the candidate host receptor can be estimated using tools such as AutoDock Vina or free energy perturbation methods [<a href="#ref-24">24</a>, <a href="#ref-28">28</a>]. PLM predictions can prioritize which mutations to test in silico, reducing the search space [<a href="#ref-29">29</a>]. For example, if a PLM predicts that a bat coronavirus spike protein can bind human ACE2, MD simulations can quantify the stability of the complex and identify key interface residues [<a href="#ref-4">4</a>]. This integrative approach combines the speed of sequence-based prediction with the physical realism of atomistic simulations [<a href="#ref-2">2</a>, <a href="#ref-27">27</a>]. The results can be further validated by surface plasmon resonance or pseudovirus entry assays, but those are outside the scope of purely computational methods.
Comparison with Traditional Phylogenetic and Machine Learning Approaches
Traditional phylogenetic analysis reconstructs viral evolutionary relationships and may infer host ancestry through tree topology [<a href="#ref-3">3</a>, <a href="#ref-5">5</a>]. However, recombination and convergent evolution can obscure these signals, and phylogenetic methods do not directly capture functional constraints at the molecular level [<a href="#ref-1">1</a>]. Classical machine learning approaches that use hand-crafted features (e.g., amino acid composition, codon usage bias, or epitope sequence motifs) have been applied to predict host tropism [<a href="#ref-12">12</a>, <a href="#ref-13">13</a>]. These models often achieve moderate accuracy but struggle to generalize to novel viruses because the features are not learned end-to-end [<a href="#ref-22">22</a>]. In contrast, PLMs automatically learn relevant features from the sequence itself, capturing long-range interactions and structural information without manual feature engineering [<a href="#ref-6">6</a>, <a href="#ref-30">30</a>]. Benchmark studies have shown that PLM-based classifiers outperform both phylogenetic and feature-based models on tasks such as distinguishing avian from human influenza strains and predicting coronavirus host origin [<a href="#ref-14">14</a>, <a href="#ref-20">20</a>]. A comparison is presented in Table 1.
| Method | Input | Feature Engineering | Generalizability | Computational Cost |
|---|---|---|---|---|
| Phylogenetic (e.g., ML tree) | Aligned sequence | Manual alignment | Moderate | High |
| Classical ML (e.g., random forest) | Hand-crafted features | Manual | Low | Low |
| Protein Language Model (PLM) | Raw sequence | Automatic (embedding) | High | Moderate |
Table 1. Comparison of traditional and PLM-based approaches for predicting viral host tropism.
Integration with 3D Protein Viewer
The residue-level attention weights produced by PLMs can be mapped onto three-dimensional structures of viral glycoproteins, such as those generated by AlphaFold2 or experimentally determined [<a href="#ref-7">7</a>]. Interactive visualization allows researchers to inspect which regions of the spike protein are most predictive of host tropism. For example, if the model attends strongly to residues in the receptor-binding domain (RBD) of a coronavirus spike, those residues can be highlighted in a 3D viewer, facilitating hypothesis generation about binding interface changes [<a href="#ref-4">4</a>]. This integration bridges the gap between abstract sequence embeddings and tangible structural biology, enabling a more intuitive understanding of zoonotic risk [<a href="#ref-8">8</a>, <a href="#ref-15">15</a>].
Limitations and Future Directions
Despite their promise, PLMs have limitations. Training data may be biased toward well-studied viruses and hosts, reducing performance for less-characterized pathogens [<a href="#ref-6">6</a>, <a href="#ref-9">9</a>]. The embeddings are not inherently interpretable, although attention weights provide some insight [<a href="#ref-16">16</a>]. Computational cost of running large transformers can be significant [<a href="#ref-17">17</a>]. Sequence-based predictions may not capture post-translational modifications or glycan shielding that affect receptor binding [<a href="#ref-19">19</a>, <a href="#ref-21">21</a>]. Future developments will likely integrate multiple data modalities, including structure, metadata, and host gene expression, into unified foundation models [<a href="#ref-13">13</a>, <a href="#ref-23">23</a>]. Hybrid approaches that combine PLMs with physics-based simulations are also promising for improving prediction accuracy [<a href="#ref-14">14</a>, <a href="#ref-24">24</a>]. Finally, rigorous prospective validation using experimental data from emerging viruses is needed to confirm the reliability of these tools for veterinary surveillance [<a href="#ref-22">22</a>].
Conclusion
Protein language models represent a significant advance in computational virology for predicting viral host tropism and zoonotic potential. By leveraging deep contextual embeddings from viral glycoprotein sequences, these models can rapidly assess the risk of cross-species transmission. Integration with receptor-binding dynamics simulations and 3D structural visualization provides a comprehensive framework for risk assessment. As the availability of viral sequence data grows and PLM architectures continue to evolve, these tools will become increasingly valuable for veterinary diagnostics and early warning of zoonotic spillover [<a href="#ref-9">9</a>, <a href="#ref-7">7</a>, <a href="#ref-1">1</a>, <a href="#ref-2">2</a>, <a href="#ref-16">16</a>, <a href="#ref-28">28</a>, <a href="#ref-17">17</a>, <a href="#ref-4">4</a>, <a href="#ref-3">3</a>, <a href="#ref-18">18</a>, <a href="#ref-30">30</a>, <a href="#ref-8">8</a>, <a href="#ref-11">11</a>, <a href="#ref-5">5</a>, <a href="#ref-29">29</a>, <a href="#ref-19">19</a>, <a href="#ref-20">20</a>, <a href="#ref-23">23</a>, <a href="#ref-27">27</a>, <a href="#ref-21">21</a>, <a href="#ref-12">12</a>, <a href="#ref-13">13</a>, <a href="#ref-24">24</a>, <a href="#ref-14">14</a>, <a href="#ref-25">25</a>, <a href="#ref-22">22</a>, <a href="#ref-26">26</a>, <a href="#ref-6">6</a>, 31, 32, 33, 34, 35].
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices