AlphaFold Database Adds Viral Protein Complexes: NVIDIA, EMBL-EBI and Glasgow
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- On 24 September 2026, EMBL-EBI, Google DeepMind, NVIDIA and partners added roughly 8,000 viral dimers to the AlphaFold Protein Structure Database, covering 2,812 viral proteomes from 23 virus families [1].
- The team analysed 41,774 viral proteins, predicted 40,746 homodimers and about 1.7 million heterodimers, and accepted 2,749 homodimers and 5,279 heterodimers after quality filtering [1].
- According to NVIDIA, about 30% of the protein interactions added have structures not previously documented in the Protein Data Bank [3].
- The earlier March 2026 release covered about 31 million complexes across 4,777 proteomes and represented 17 million GPU hours, but included no viruses [4][5].
- The AlphaFold Database now holds more than 260 million protein and protein complex predictions and has about 3.4 million users in 190 countries [1].
- Glycans, assemblies larger than dimers, and trimer predictions for coronavirus spikes and the HIV envelope protein are all missing from this release [1][2].
On 24 September 2026, the AlphaFold Protein Structure Database gained its first set of AlphaFold viral protein complexes. The release covers 2,812 viral proteomes drawn from 23 virus families, every one of which contains members that infect humans. After filtering, about 8,000 viral dimers entered the database, and the whole set sits behind a new Pandemic Preparedness Portal on the AlphaFold homepage [1].
Why it matters: until this year, almost everything in the database was a single protein chain. Viruses rarely work that way. Polymerases, proteases, capsid building blocks and the viral proteins that grab host factors mostly operate as pairs or larger assemblies, and those interfaces are exactly where antibodies and drugs tend to land. The new release does not solve viral structural biology. It gives the field a large, open, machine-readable starting point for it.
Why single-chain structures were never enough for virology
The AlphaFold Protein Structure Database launched in 2021, run by EMBL's European Bioinformatics Institute with Google DeepMind. Before the complexes work, it held predictions for over 200 million proteins, all of them single chains, or monomers. The AlphaFold team received the 2024 Nobel Prize in Chemistry for the underlying method.
A monomer prediction answers a narrow question: given this sequence, what shape does this one chain fold into? For a great many viral proteins, that is only half the story. Two copies of the same protein can pair up into a homodimer. Two different proteins can pair into a heterodimer. Both arrangements create new surfaces that do not exist on either chain alone, and those surfaces are often the functional ones. If you want to understand how a viral polymerase assembles, or where a neutralising antibody binds, you need the quaternary structure. The fold of one chain will not show it.
There is also a practical reason this matters for anyone doing bench work. Interface residues are attractive mutagenesis targets. Interface-blocking peptides and small molecules are designed against them. Cryo-EM and crystallography projects get prioritised around them. A database of confident dimer predictions turns a slow literature search into a query.
Step one: Viro3D brought 4,400 viruses into the AlphaFold Database
The groundwork came from Glasgow. Viro3D, from the MRC-University of Glasgow Centre for Virus Research, was published in Molecular Systems Biology on 16 September 2025 by Ulad Litvin, Spyros Lytras, Alexander Jack, David L Robertson, Joseph Hughes and Joe Grove [6]. The team predicted structures for more than 85,000 proteins from more than 4,400 viruses, running two tools side by side: ColabFold, which is AlphaFold2-based, produced 85,162 predictions, and ESMFold, a protein language model, produced 84,964 [6].

Figure 1. How Viro3D was built and what it found, based on Litvin et al. 2025 [6]. Schematic, AI-assisted illustration.
That expands structural coverage of viral proteins about 30 times compared with experimental structures, with about 27 million residues modelled and 19,067 structural clusters [6]. The novelty finding is the striking part. 82% of viral protein clusters had no detectable similarity to proteins from cellular organisms [6]. In other words, most of what viruses build does not look like anything in a human cell, which is part of why antiviral drug development is hard.
Viro3D also turned up 251 class I fusion glycoproteins, including previously unrecognised ones in herpesviruses and baculoviruses. Analysis in the paper suggests the coronavirus spike glycoprotein may have originated through genetic exchange with an ancestral herpesvirus [6]. That is the kind of claim that would have been very hard to assemble without a broad comparative structure set.
| Viro3D at a glance | Value |
|---|---|
| Proteins predicted | More than 85,000 |
| Viruses covered | More than 4,400 |
| ColabFold predictions | 85,162 |
| ESMFold predictions | 84,964 |
| Coverage gain vs experimental structures | About 30 times |
| Residues modelled | About 27 million |
| Structural clusters | 19,067 |
| Clusters with no detectable similarity to cellular proteins | 82% |
| Class I fusion glycoproteins found | 251 |
On 25 February 2026 the University of Glasgow announced that Viro3D predictions, covering 85,000+ proteins from 4,400+ viruses, were now accessible through the AlphaFold Database, as part of the AFDB opening to community datasets [7]. Joe Grove framed it plainly at the time: "The AlphaFold revolution is having a transformative impact on molecular biology, providing new opportunities for fundamental discovery and accelerating the development of therapies."
The limitation carried over with it. Viro3D, like every monomer database, holds single chains only. Most viral proteins work with partners.
Step two: NVIDIA and Seoul National University scale AlphaFold to 31 million complexes
The first protein complexes landed in the AFDB on 16 March 2026, from EMBL-EBI, Google DeepMind, NVIDIA and Seoul National University [4]. The preprint, "AlphaFold Database expands to proteome-scale quaternary structures", lists first authors Yewon Han, Maxim I. Tsenkov and Niccolo A. E. Venanzi, with senior authors including Christian Dallago, Milot Mirdita and Jennifer Fleming, and affiliations spanning Seoul National University, EMBL-EBI, NVIDIA, Google DeepMind and Duke University [5].
The scale: about 31 million complexes predicted, made up of 23.4 million homodimers from 4,777 proteomes and 7.6 million heterodimer candidates drawn from STRING physical interactions. The proteomes included 16 model organisms, human among them, plus 30 WHO global health priority proteomes. No viruses in this release [4][5].
Of those, 1.7 to 1.8 million high-confidence complexes went into the AFDB. EMBL says 1.7 million high-confidence homodimers; the preprint says 1.8 million high-confidence complexes [4][5]. By a May 2026 update, about 80,000 high-confidence heterodimers were also in. The lower-confidence sets stayed available for bulk download: 18 million homodimers and 8.1 million heterodimers [4].
EMBL states the release represents 17 million GPU hours [4].
How the pipeline works
The compute story explains why this became possible in 2026 and not earlier. Multiple sequence alignments were built with MMseqs2-GPU, developed by Martin Steinegger's group at Seoul National University with NVIDIA, against UniRef100. The pipeline kept only the best hit per taxon, which avoids diluting alignments with paralogs. Structures were then predicted with AlphaFold-Multimer, run through ColabFold and through OpenFold accelerated with NVIDIA TensorRT and cuEquivariance, on an NVIDIA DGX H100 SuperPOD [5].
flowchart TD
A["Proteome input"] --> B["MMseqs2-GPU MSA against UniRef100"]
B --> C["Keep best hit per taxon"]
C --> D["AlphaFold-Multimer via ColabFold"]
C --> E["OpenFold with TensorRT and cuEquivariance"]
D --> F["Scoring ipSAE pLDDT clashes"]
E --> F
F --> G["High confidence into AFDB"]
F --> H["Lower confidence to bulk download"]
Martin Steinegger put the biological motivation simply: "Biology really happens when things come together." [8]
How confident is "high confidence"?
For high-confidence homodimers, three filters had to pass together: ipSAE_min at least 0.6, average pLDDT at least 70, and backbone clash score at most 10. pLDDT is the per-residue confidence score on a 0 to 100 scale. ipSAE is a newer interface score that concentrates on the residues actually sitting at the interface, which makes it less inflated by disordered regions than older interface metrics.
| Display tier | ipSAE_min threshold | Entries |
|---|---|---|
| Very high confidence | 0.8 and above | 972,625 |
| Confident | 0.7 to below 0.8 | 438,879 |
| Low confidence | 0.6 to below 0.7 | 342,738 |
Validation used 1,968 homodimers from the Protein Data Bank released after 30 September 2021, plus 2,211 monomeric negative controls. At the ipSAE_min 0.6 threshold, precision was 0.859, recall 0.655 and F1 0.744 [5]. That is a useful number to hold onto. Roughly one in seven confident-looking homodimers is not a real dimer by this benchmark, and about a third of true dimers fall below the threshold.
The authors were direct about the rest. Complex prediction has not reached experimental accuracy and shows biases. Heterodimer confidence calibration was insufficient, so 56,959 heterodimers were labelled "tentatively high-confidence". High-confidence heterodimers skew toward pairs of similar length and higher sequence identity. Prediction success was more than 3-fold higher in prokaryotes than eukaryotes [5].
Jo McEntyre of EMBL-EBI, announcing the March release, said: "By making this foundational dataset openly available to the world, we're inviting researchers to test, refine, and build on it to drive the next wave of biological discoveries." [4]
Step three: AlphaFold viral protein complexes and the Pandemic Preparedness Portal
The September 2026 release is where viruses finally enter. The partners were EMBL-EBI, Google DeepMind, NVIDIA, Seoul National University, Sungkyunkwan University in South Korea, the MRC-University of Glasgow Centre for Virus Research, the Swiss Institute of Bioinformatics, the Coalition for Epidemic Preparedness Innovations and Lund University in Sweden [1].
The scope: 2,812 viral proteomes, over 2,800 viruses, from 23 virus families, every family containing members that infect humans. 41,774 viral proteins were analysed. The team predicted 40,746 homodimers and about 1.7 million heterodimers. After quality filtering, 2,749 homodimers and 5,279 heterodimers were accurate enough to include, roughly 8,000 viral dimers in the AFDB [1].
| September 2026 release | Value |
|---|---|
| Viral proteomes | 2,812 |
| Virus families | 23 |
| Viral proteins analysed | 41,774 |
| Homodimers predicted | 40,746 |
| Heterodimers predicted | About 1.7 million |
| Homodimers accepted | 2,749 |
| Heterodimers accepted | 5,279 |
| Total viral dimers in AFDB | About 8,000 |
| Interactions with structures new to the PDB | About 30% |

Figure 2. Homodimers versus heterodimers, and the three filters a predicted complex must pass to be labelled high confidence [5]. Schematic, AI-assisted illustration.
According to NVIDIA, about 30% of the protein interactions added have structures not previously documented in the Protein Data Bank [3]. That figure is the clearest signal of what this adds. These are not re-annotations of known structures. They are models of interfaces with no deposited experimental structure.
The choice of viruses was guided by the UK Health Security Agency priority pathogen tool, and coverage includes Picornaviridae and mpox. Viruses named in the coverage include mpox, measles, hepatitis B, Zika and dengue [1]. Access is open, through a Pandemic Preparedness Portal on the AFDB homepage at alphafold.ebi.ac.uk [1]. The timing was deliberate: the release coincided with the UN General Assembly High-level Meeting on Pandemic Prevention, Preparedness and Response on 25 September 2026, and it was framed around the 100 Days Mission, the goal of having vaccines within 100 days of an outbreak [1][3].
What NVIDIA contributed
The models are AlphaFold2 and AlphaFold-Multimer. Inference was optimised with the NVIDIA BioNeMo Inference Runtime, and NVIDIA openly released the GPU-accelerated BioNeMo Structure Prediction Pipeline used to produce the set [3]. The company's framing of the goal came from Chris Dallago: "We're enabling biologists and the AI community to investigate protein interactions, not just as single molecules but as complexes." [3]
Who did what
EMBL-EBI runs the AlphaFold Protein Structure Database with Google DeepMind and hosted the release, with Jo McEntyre, Interim Director, saying: "Making these data open is critical for understanding viral diagnostics and developing treatments and vaccines." [1] Sameer Velankar of EMBL-EBI was part of both complex releases [1][4].
Google DeepMind built AlphaFold2 and AlphaFold-Multimer, the models behind every prediction. NVIDIA supplied the accelerated inference stack through the BioNeMo runtime and released the pipeline. Seoul National University brought the MMseqs2-GPU alignment work that made the March release feasible, and Sungkyunkwan University joins it as a Korean partner on the viral release. The Glasgow centre is where Viro3D was built, and Joe Grove's group supplies the virology. The Swiss Institute of Bioinformatics, CEPI and Lund University complete the partner list [1]. The press material does not break down each partner's exact task, so treat this as the broad division of labour and not a line-by-line credit.
Joe Grove, Professor of Molecular Virology at the Glasgow centre, described the underlying problem to Nature: "Many viral proteins do not act individually, they act in concert with partners." [2] He also described the intent behind the work in the NVIDIA blog: "What we're trying to do is stockpile some of that knowledge ahead of time." [3]
Limitations: glycans, trimers and mutations
The limitations are real, and the team stated them clearly.
Predictions lack the sugar chains, or glycans, that coat many viral surface proteins and help them evade antibodies. A model without glycans shows the protein backbone and side chains but not the shield that antibodies actually encounter.
Many viral proteins act in assemblies larger than dimers. The release covers dimers only, so anything that needs a trimer, tetramer or higher-order capsid arrangement is out of scope for now.
The most consequential gap is trimers. The spike proteins of SARS-CoV-2 and other coronaviruses, and the HIV envelope protein, are trimers, and trimer predictions were not accurate enough to include [1][2]. Given how much pandemic preparedness work centres on coronavirus spikes, that is a real hole. Anyone working on bat coronaviruses or related spike biology should read this release as context. It does not replace trimer structures.
The structures also do not predict the effect of mutations, host-pathogen interactions, or changes in transmissibility or virulence. Experimental confirmation is needed [1][2]. Grove made the point sharply in the EMBL release: "A protein complex structure alone doesn't tell us what happens when a virus mutates." [1]
Anyone who has worked on viral variants at the bench will recognise the problem. A predicted interface tells you where to look. It does not tell you what a single amino acid substitution at that interface does to binding, assembly or fitness. That still requires the wet lab.
How virologists and drug hunters can use the data now
The practical value depends on treating these models as hypotheses with confidence scores attached. They are not answers. Five concrete uses:
- Pick candidate interfaces for mutagenesis. Take a high-confidence homodimer, identify the residues at the interface, and mutate them to test whether the interaction matters in cells. The ipSAE_min tier tells you how much to trust the interface in the first place.
- Design interface-blocking peptides or small molecules. Interface surfaces are often shallow and hydrophobic, which makes them hard drug targets, but they are still starting points for peptide design and for fragment screening.
- Check whether a new variant sits at an interface. If a mutation of concern falls on a predicted dimer interface, that is a reason to prioritise it for functional work. If it sits in a disordered loop, it probably is not about assembly.
- Prioritise targets for cryo-EM or crystallography. A confident model narrows the list. A low-confidence model tells you the sample may be hard to prepare, which is itself useful information.
- Train and benchmark new AI models on the bulk data. The 18 million lower-confidence homodimers and 8.1 million heterodimers are available for bulk download, which is a large training set for anyone working on interface prediction [4].
How to access it, step by step. Go to alphafold.ebi.ac.uk and look for the Pandemic Preparedness Portal on the homepage. Find your protein or complex of interest, then check two things before you use the model: the average pLDDT, which tells you how well the individual chains are modelled, and the interface confidence score, which tells you how much to trust the contact between them. Anything in the very high confidence tier (ipSAE_min 0.8 and above) is a reasonable basis for planning experiments. Anything in the low confidence tier (0.6 to below 0.7) is a hint, not a result. For machine learning work, use the bulk download instead of the portal. To inspect a downloaded PDB or mmCIF file without installing PyMOL or ChimeraX, open it in the protein structure viewer on this site.
If the terminology here is unfamiliar, the broader framing of how structure prediction fits into data-driven biology is covered in this piece on computational biology.
The timeline in one picture
flowchart TD
A["2021 AFDB launch"] --> B["Sept 2025 Viro3D paper"]
B --> C["Feb 2026 Viro3D in AFDB"]
C --> D["Mar 2026 31M complexes"]
D --> E["May 2026 heterodimer update"]
E --> F["Sept 2026 viral complexes portal"]
Frequently Asked Questions
What are AlphaFold viral protein complexes?
They are predicted three-dimensional structures of two viral protein chains bound together, produced by AlphaFold2 and AlphaFold-Multimer and stored in the AlphaFold Protein Structure Database. The September 2026 release added about 8,000 of them, covering 2,812 viral proteomes from 23 virus families [1]. They are computational models, not experimental structures.
Does the AlphaFold Database now include viruses?
Yes. The March 2026 release of 31 million complexes deliberately excluded viruses, but the September 2026 release added 2,749 homodimers and 5,279 heterodimers from viral proteomes [1][4]. Viro3D predictions for more than 85,000 viral proteins had already been accessible through the database since February 2026 [7].
What is the difference between a homodimer and a heterodimer in the AlphaFold Database?
A homodimer is two copies of the same protein. A heterodimer is two different proteins. The March 2026 release predicted 23.4 million homodimers and 7.6 million heterodimer candidates, and the heterodimer set was drawn from STRING physical interactions [5]. Heterodimer confidence calibration was weaker, which is why 56,959 were labelled "tentatively high-confidence" [5].
How accurate are the AlphaFold viral complex predictions?
The team validated against 1,968 homodimers from the Protein Data Bank released after 30 September 2021 and 2,211 monomeric negative controls. At the ipSAE_min 0.6 threshold, precision was 0.859, recall 0.655 and F1 0.744 [5]. The authors state that complex prediction has not reached experimental accuracy and shows biases, including more than 3-fold better success in prokaryotes than eukaryotes.
Why are glycans and trimers missing from this release?
Glycans are not modelled, so the sugar coating that shields many viral surface proteins from antibodies is absent. Trimer predictions, including coronavirus spikes and the HIV envelope protein, were not accurate enough to include [1][2]. Many viral proteins also act in assemblies larger than dimers, which are outside the scope of this release.
How do I access the viral protein complexes in the AlphaFold Database?
Go to alphafold.ebi.ac.uk and use the Pandemic Preparedness Portal on the homepage. Check the average pLDDT and the interface confidence score before using any model. For large-scale work, the lower-confidence sets are available for bulk download: 18 million homodimers and 8.1 million heterodimers [4].
Can AlphaFold predict how a viral mutation changes transmissibility?
No. The structures do not predict the effect of mutations, host-pathogen interactions, or changes in transmissibility or virulence, and experimental confirmation is needed [1][2]. Joe Grove put it directly: "A protein complex structure alone doesn't tell us what happens when a virus mutates." [1]
Related Articles
- AlphaGenome explained: DeepMind's DNA model and the 9-billion-variant atlas
- Quaternary structure of proteins
- Bat coronaviruses: veterinary and One Health reference
- What is computational biology?
References
- AlphaFold Database adds viral protein complexes to support pandemic preparedness. EMBL, 24 September 2026.
- Callaway E. AlphaFold 'goes viral': database adds protein complexes of common viruses. Nature, September 2026.
- How Open Science Can Help Researchers Prepare for the Next Pandemic. NVIDIA Blog, 24 September 2026.
- Millions of protein complexes added to AlphaFold Database shed light on how proteins interact. EMBL, 16 March 2026 (updated 19 May 2026).
- Han Y, Tsenkov MI, Venanzi NAE, et al. AlphaFold Database expands to proteome-scale quaternary structures. bioRxiv, 2026.
- Litvin U, Lytras S, Jack A, Robertson DL, Hughes J, Grove J. Viro3D: a comprehensive database of virus protein structure predictions. Molecular Systems Biology, 2025.
- Viro3D structure predictions now available through the AlphaFold Database. University of Glasgow Centre for Virus Research, 25 February 2026.
- AlphaFold Can Now Predict Protein Complex Structures at Scale. The Scientist, March 2026.
- Litvin U, et al. Viro3D (PubMed record). PubMed 40958060