Machine Learning for Predicting Host-Virus Protein-Protein Interaction Networks
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Machine learning (ML) models are crucial for predicting host-virus protein-protein interactions (PPIs), overcoming the limitations of experimental methods in terms of cost, time, and scalability. These interactions are fundamental to viral entry, replication, immune evasion, and pathogenesis, making their prediction vital for understanding disease mechanisms.
- Advanced feature engineering, including amino acid composition, evolutionary information (PSSMs, co-evolution), and deep protein sequence embeddings (e.g., word2vec, doc2vec), significantly enhances the predictive power of ML models for host-virus PPIs. These features capture complex molecular characteristics that drive binding specificity.
- Deep learning architectures, particularly Convolutional Neural Networks (CNNs) and Long Short-Term Memory (LSTM) networks, along with Graph Neural Networks (GNNs), represent the state-of-the-art. CNNs excel at identifying local motifs, LSTMs capture long-range dependencies, and GNNs integrate network topology, leading to improved prediction accuracy.
- Transfer learning is a critical strategy for veterinary virology, enabling the application of models trained on well-studied human-virus interactions to predict interactions for novel or under-characterized animal pathogens. This approach is essential given the scarcity of experimental data for many veterinary viruses.
- Applications in veterinary virology include predicting host range and zoonotic potential, identifying novel antiviral therapeutic targets by revealing host dependencies, understanding pathogenesis by mapping hijacked host pathways, and aiding vaccine design by pinpointing critical interaction interfaces.
- Challenges remain in addressing data imbalance through robust negative sampling strategies and improving model generalization across diverse virus families. Future directions involve integrating structural information from tools like AlphaFold2 and developing more interpretable ML models with enhanced veterinary-specific datasets.
Introduction
The molecular interface between a viral pathogen and its host cell is fundamentally governed by protein-protein interactions (PPIs) [<a href="#ref-1">1</a>, <a href="#ref-2">2</a>]. These interactions mediate viral entry, replication, immune evasion, and pathogenesis [<a href="#ref-3">3</a>, <a href="#ref-4">4</a>]. Experimental identification of host-virus PPIs, such as yeast two-hybrid screens, co-immunoprecipitation, and affinity purification mass spectrometry, remains time-consuming, costly, and not scalable to the vast combinatorial space of potential interactions [<a href="#ref-5">5</a>, <a href="#ref-6">6</a>]. Computational methods, particularly machine learning (ML) approaches, have emerged as powerful tools to predict host-virus PPIs with high throughput and reasonable accuracy [<a href="#ref-7">7</a>, <a href="#ref-8">8</a>]. This review provides an exhaustive, publication-grade examination of ML methodologies for predicting host-virus PPI networks, with a focus on graph neural networks, protein co-evolution, sequence patterns, and structural pocket mapping. The discussion is framed within the context of veterinary virology, drawing parallels to human systems where relevant but emphasizing applications to animal pathogens.
Biological Basis of Host-Virus Protein-Protein Interactions
Host-virus PPIs occur when viral proteins bind to specific host proteins to hijack cellular machinery [<a href="#ref-9">9</a>, <a href="#ref-10">10</a>]. The binding interface typically involves complementary surface patches characterized by specific amino acid compositions, electrostatic potentials, and hydrophobic contacts [<a href="#ref-2">2</a>, <a href="#ref-11">11</a>]. Viral proteins often mimic host interaction motifs to subvert normal signaling pathways [<a href="#ref-12">12</a>, <a href="#ref-13">13</a>]. For example, influenza A virus NS1 protein binds to host CPSF30 to inhibit host mRNA processing [<a href="#ref-14">14</a>, <a href="#ref-15">15</a>]. Similarly, coronavirus spike proteins engage host ACE2 receptors to initiate membrane fusion [<a href="#ref-16">16</a>, <a href="#ref-17">17</a>]. Understanding these interactions at the molecular level is critical for predicting host range, zoonotic potential, and therapeutic targets [<a href="#ref-18">18</a>, <a href="#ref-19">19</a>].
The structural determinants of binding can be captured through features such as amino acid composition, dipeptide composition, conjoint triad, pseudo amino acid composition, and autocorrelation descriptors [<a href="#ref-20">20</a>, <a href="#ref-21">21</a>]. Additionally, evolutionary information encoded in position-specific scoring matrices (PSSMs) and co-evolutionary signals between interacting protein families provide rich predictive features [<a href="#ref-22">22</a>, <a href="#ref-23">23</a>]. The three-dimensional geometry of binding pockets, including solvent accessibility, secondary structure propensities, and residue depth, further refines interaction predictions [<a href="#ref-2">2</a>, <a href="#ref-24">24</a>].
Machine Learning Approaches for Host-Virus PPI Prediction
Feature Engineering and Representation
Transforming protein sequences into numerical feature vectors is a critical preprocessing step [<a href="#ref-5">5</a>, <a href="#ref-25">25</a>]. Two primary paradigms exist: Independent Protein Feature (IPF) extraction, where host and viral proteins are encoded separately, and Merged Protein Feature (MPF) extraction, where sequences are concatenated before encoding [<a href="#ref-5">5</a>]. An Extended Protein Feature (EPF) method that combines both approaches has shown improved performance with traditional ML classifiers such as Support Vector Machine (SVM), Logistic Regression, and Multilayer Perceptron [<a href="#ref-5">5</a>].
Common sequence-based features include:
- Amino acid composition (AAC) and pseudo amino acid composition (PAAC) [<a href="#ref-7">7</a>, <a href="#ref-16">16</a>]
- Dipeptide composition (DC) and conjoint triad (CT) [<a href="#ref-22">22</a>, <a href="#ref-26">26</a>]
- Moran autocorrelation and normalized Moreau-Broto autocorrelation [<a href="#ref-22">22</a>]
- Tripeptide features derived from reduced amino acid alphabets [<a href="#ref-12">12</a>]
- Gene ontology (GO) terms and natural language processing (NLP) embeddings such as bag-of-words, TF-IDF, and doc2vec [<a href="#ref-7">7</a>, <a href="#ref-14">14</a>, <a href="#ref-21">21</a>]
Deep protein sequence embeddings, inspired by NLP, have revolutionized feature representation [<a href="#ref-27">27</a>, <a href="#ref-28">28</a>]. Methods like word2vec, doc2vec, and Byte Pair Encoding (BPE) convert amino acid sequences into dense vector representations that capture contextual and semantic relationships [<a href="#ref-20">20</a>, <a href="#ref-21">21</a>]. The Siamese Tailored deep sequence Embedding of Proteins (STEP) approach integrates these embeddings into a Siamese neural network architecture to predict virus-host PPIs [<a href="#ref-27">27</a>].
Classical Machine Learning Classifiers
Several supervised learning algorithms have been benchmarked for host-virus PPI prediction [<a href="#ref-6">6</a>, <a href="#ref-9">9</a>]. Random Forest (RF) and Support Vector Machine (SVM) are among the most frequently used, with RF often achieving high accuracy due to its ensemble nature [<a href="#ref-6">6</a>, <a href="#ref-26">26</a>]. Deep Forest, an ensemble of cascade forest models, has demonstrated superior predictive accuracy compared to RF and SVM in some studies [<a href="#ref-6">6</a>]. Other classifiers include Logistic Regression, Naive Bayes, K-Nearest Neighbors, and eXtreme Gradient Boosting (XGBoost) [<a href="#ref-6">6</a>, <a href="#ref-22">22</a>, <a href="#ref-29">29</a>]. Feature selection methods, such as correlation coefficient-based filtering and Max-Min Parents and Children (MMPC), reduce dimensionality while maintaining predictive performance [<a href="#ref-7">7</a>, <a href="#ref-12">12</a>].
Deep Learning Architectures
Deep neural networks (DNNs) have become the state-of-the-art for host-virus PPI prediction due to their ability to learn hierarchical features from raw sequence data [<a href="#ref-4">4</a>, <a href="#ref-24">24</a>]. Convolutional neural networks (CNNs) extract local, position-dependent motifs, while long short-term memory (LSTM) networks capture long-range dependencies [<a href="#ref-4">4</a>, <a href="#ref-21">21</a>, <a href="#ref-25">25</a>]. Hybrid architectures combining CNN and LSTM layers have been proposed to leverage both local and sequential information [<a href="#ref-4">4</a>, <a href="#ref-25">25</a>]. The LSTM-PHV model, using word2vec embeddings, achieved an AUC of 0.976 and accuracy of 98.4% on human-virus datasets [<a href="#ref-21">21</a>].
Multi-scale CNNs and Siamese CNN architectures have been employed to compare pairs of protein sequences directly [<a href="#ref-17">17</a>, <a href="#ref-30">30</a>]. The Siamese network processes host and viral proteins through identical subnetworks, learning a similarity metric that indicates interaction likelihood [<a href="#ref-17">17</a>, <a href="#ref-27">27</a>, <a href="#ref-30">30</a>]. Transfer learning strategies, including frozen and fine-tuning approaches, enable knowledge transfer from well-characterized virus-host systems to novel or understudied viruses [<a href="#ref-11">11</a>, <a href="#ref-17">17</a>, <a href="#ref-31">31</a>].
Graph Neural Networks and Network Topology
Graph neural networks (GNNs) incorporate topological information from known protein-protein interaction networks to generate enriched protein embeddings [<a href="#ref-20">20</a>, <a href="#ref-32">32</a>]. The GraphSAGE model, a variant of graph convolutional networks, aggregates features from neighboring nodes in the host PPI network to produce hybrid embeddings that combine sequence-derived and network-derived features [<a href="#ref-20">20</a>]. This approach improved AUC scores by 3-23% over sequence-only methods [<a href="#ref-20">20</a>]. Incorporating viral molecular mimicry and network centrality measures (e.g., degree, betweenness, closeness) further enhances prediction accuracy [<a href="#ref-26">26</a>, <a href="#ref-32">32</a>].
Transfer Learning and Multi-Task Learning
Given the scarcity of experimentally validated PPIs for many veterinary viruses, transfer learning is particularly valuable [<a href="#ref-11">11</a>, <a href="#ref-31">31</a>]. Models pre-trained on large human-virus interaction datasets can be fine-tuned on smaller animal virus datasets [<a href="#ref-11">11</a>, <a href="#ref-17">17</a>]. The DeepVHPPI framework uses a self-attention-based transformer with transfer learning to predict interactions for novel virus sequences, including mutated variants [<a href="#ref-11">11</a>]. Probability-weighted ensemble transfer learning has been applied to HIV-human interactions, demonstrating robustness to data unavailability [<a href="#ref-31">31</a>].
Evaluation Metrics and Benchmarking
Standard evaluation metrics for binary classification include accuracy, precision, recall, F1-score, Matthews Correlation Coefficient (MCC), and area under the receiver operating characteristic curve (AUC-ROC) [<a href="#ref-6">6</a>, <a href="#ref-9">9</a>]. Cross-validation (e.g., 5-fold or 10-fold) is used to assess generalization [<a href="#ref-21">21</a>, <a href="#ref-29">29</a>]. Independent test sets from different virus families provide rigorous validation of model transferability [<a href="#ref-4">4</a>, <a href="#ref-11">11</a>]. The following table summarizes representative performance metrics from key studies:
| Model | Dataset | Accuracy | AUC | F1-score | Reference |
|---|---|---|---|---|---|
| LSTM-PHV | Human-virus | 98.4% | 0.976 | - | [<a href="#ref-21">21</a>] |
| XGBoost (IAV-human) | Influenza A-human | 96.89% | - | 96.78% | [<a href="#ref-22">22</a>] |
| Deep Forest | Human-parasite/bacteria | Highest | - | - | [<a href="#ref-6">6</a>] |
| Hybrid CNN-LSTM | Virus-host benchmark | Best | - | - | [<a href="#ref-4">4</a>, <a href="#ref-25">25</a>] |
| GraphSAGE hybrid | Human-virus | - | 3-23% better | - | [<a href="#ref-20">20</a>] |
| MMPC-DNN (GO-NLP) | SARS-CoV-2 | - | 0.878 | 0.793 | [<a href="#ref-7">7</a>] |
Workflow for Host-Virus PPI Prediction
The typical computational pipeline integrates data collection, feature extraction, model training, and validation. The following Mermaid diagram illustrates a generalized workflow:
flowchart TD
A["Experimental PPI Databases<br/>e.g., HVIDB, STRING"] --> B["Data Preprocessing<br/>Positive & Negative Sampling"]
B --> C["Feature Extraction"]
C --> D["Sequence-based Features<br/>AAC, CT, DC, PSSM, Embeddings"]
C --> E["Structural Features<br/>Solvent Accessibility, Pocket Geometry"]
C --> F["Network Topology Features<br/>Centrality, Graph Embeddings"]
D --> G["Feature Selection<br/>Correlation, MMPC, PCA"]
E --> G
F --> G
G --> H["Model Training"]
H --> I["Classical ML<br/>RF, SVM, XGBoost"]
H --> J["Deep Learning<br/>CNN, LSTM, Siamese, GNN"]
H --> K["Transfer Learning<br/>Frozen/Fine-tuning"]
I --> L["Evaluation<br/>Cross-validation, Independent Test"]
J --> L
K --> L
L --> M["Prediction of Novel PPIs"]
M --> N["Validation via GO/KEGG Enrichment"]
N --> O["3D Structural Mapping<br/>Binding Pocket Visualization"]
Structural Pocket Mapping and 3D Visualization
Predicting which residues at the binding interface contribute most to the interaction is a key downstream task [<a href="#ref-19">19</a>, <a href="#ref-27">27</a>]. Explainable AI (XAI) methods, such as attention weights and gradient-based saliency maps, identify sequence regions critical for interaction [<a href="#ref-27">27</a>]. These regions can be mapped onto three-dimensional protein structures using tools like PyMOL or ChimeraX [<a href="#ref-2">2</a>, <a href="#ref-33">33</a>]. For veterinary applications, structural models of viral proteins (e.g., feline coronavirus spike, avian influenza hemagglutinin) can be obtained from the Protein Data Bank or predicted using AlphaFold2 [<a href="#ref-24">24</a>, <a href="#ref-33">33</a>]. Binding pocket mapping highlights conserved hydrophobic patches, charged residues, and hydrogen-bonding networks that mediate host recognition [<a href="#ref-2">2</a>, <a href="#ref-19">19</a>]. The HVIDB database provides pre-computed 3D complex structures for many human-virus PPIs, serving as a template for analogous animal systems [<a href="#ref-33">33</a>].
Applications in Veterinary Virology
While most ML-based PPI predictors have been developed for human viruses, the underlying methodologies are directly transferable to veterinary pathogens [<a href="#ref-1">1</a>, <a href="#ref-3">3</a>]. Key applications include:
- Predicting host range and zoonotic potential: ML models trained on known interactions can assess whether a viral protein from an animal reservoir (e.g., bat coronavirus, avian influenza) can bind to human or livestock receptors [<a href="#ref-18">18</a>, <a href="#ref-32">32</a>]. This complements structural approaches such as those described in Predicting Viral Host Range and Zoonotic Potential Using Machine Learning on Spike Protein Structures.
- Identifying antiviral targets: Predicted PPIs between viral proteins and host factors reveal dependencies that can be targeted by therapeutics [<a href="#ref-9">9</a>, <a href="#ref-13">13</a>]. For example, interactions between hepatitis E virus and human proteins linked to hepatocellular carcinoma have been predicted using ML [<a href="#ref-13">13</a>].
- Understanding pathogenesis: Network analysis of predicted PPIs can identify host pathways hijacked by viruses, such as immune signaling or apoptosis [<a href="#ref-16">16</a>, <a href="#ref-26">26</a>]. This is relevant for diseases like African swine fever, where viral proteins modulate host defenses [<a href="#ref-18">18</a>].
- Vaccine design: Knowledge of PPI interfaces guides the design of subunit vaccines that block viral attachment [<a href="#ref-19">19</a>, <a href="#ref-22">22</a>]. The Machine Learning in Predicting Protein-Protein Interactions article provides additional context.
- Monitoring viral evolution: ML models can rapidly assess how mutations in viral proteins (e.g., influenza hemagglutinin, coronavirus spike) alter interaction profiles, aiding surveillance of emerging strains [<a href="#ref-11">11</a>, <a href="#ref-19">19</a>]. This aligns with the Deep Mutational Scanning and Machine Learning for Predicting SARS-CoV-2 Spike Protein Escape Mutations from Antibody Neutralization article.
Databases such as HVIDB [<a href="#ref-33">33</a>] and Viruses.STRING [<a href="#ref-29">29</a>] provide extensive collections of experimentally verified and predicted host-virus PPIs that can be leveraged for veterinary species through orthology mapping.
Challenges and Future Directions
Despite significant progress, several challenges remain:
- Data imbalance and negative sampling: Experimentally validated non-interacting pairs are rarely reported, necessitating computational negative sampling strategies that may introduce bias [<a href="#ref-10">10</a>, <a href="#ref-31">31</a>]. Co-localization-based negative sampling is more reliable than random sampling [<a href="#ref-31">31</a>].
- Generalization across virus families: Models trained on one virus family often perform poorly on distantly related viruses due to differences in interaction mechanisms [<a href="#ref-11">11</a>, <a href="#ref-17">17</a>]. Transfer learning partially addresses this but requires careful domain adaptation [<a href="#ref-30">30</a>].
- Incorporation of structural information: Most sequence-based methods ignore 3D conformation, which is critical for binding specificity [<a href="#ref-2">2</a>, <a href="#ref-24">24</a>]. The advent of AlphaFold2 and other structure prediction tools enables integration of predicted structures into PPI models [<a href="#ref-24">24</a>, <a href="#ref-27">27</a>].
- Interpretability: Deep learning models are often black boxes; XAI methods are needed to identify biologically meaningful interaction determinants [<a href="#ref-27">27</a>].
- Veterinary-specific resources: Public databases are heavily biased toward human pathogens. Curating high-quality PPI datasets for livestock, poultry, and companion animals is essential for advancing veterinary applications [<a href="#ref-1">1</a>, <a href="#ref-3">3</a>].
Future directions include the development of foundation models pre-trained on large protein sequence corpora and fine-tuned for host-virus interaction prediction [<a href="#ref-24">24</a>, <a href="#ref-27">27</a>]. Graph neural networks that jointly model host and viral protein interaction networks will likely become standard [<a href="#ref-20">20</a>, <a href="#ref-32">32</a>]. Integration with molecular dynamics simulations can validate predicted binding interfaces and estimate binding affinities [<a href="#ref-2">2</a>, <a href="#ref-16">16</a>].
Conclusion
Machine learning has become an indispensable tool for predicting host-virus protein-protein interaction networks. From classical classifiers to deep learning architectures and graph neural networks, these methods enable high-throughput, cost-effective identification of molecular interactions that underpin viral infection. Feature engineering leveraging sequence, structural, and network topology information continues to improve predictive accuracy. Transfer learning and explainable AI enhance model applicability and interpretability. While most current work focuses on human viruses, the methodologies are directly applicable to veterinary virology, offering opportunities to predict host range, identify therapeutic targets, and monitor viral evolution. Continued development of veterinary-specific databases and integration with structural biology will further advance this field.