Virus Structure
Virus structure is the physical arrangement of components that form a virus particle, including a nucleic acid genome (DNA or RNA) contained within a protein shell called a capsid, which is sometimes surrounded by a lipid envelope derived from a host membrane. This guide is for life sciences researchers, bioinformatics analysts, and students who need a source bounded framework for interpreting viral architecture from experimental data and public repositories. Use this guide when you are analyzing viral sequences, designing structural studies, or evaluating published virus models. NCBI Bookshelf provides a foundational reference for viral classification and morphology. EMBL-EBI Training offers structured modules on structural bioinformatics that support the concepts discussed here.
At a Glance: Key Attributes of Virus Structure
| Attribute | Description | Common Examples |
|---|---|---|
| Genome Type | DNA or RNA, single or double stranded, segmented or linear | SARS CoV 2 (ssRNA+), Adenovirus (dsDNA) |
| Capsid Symmetry | Icosahedral, helical, or complex | Icosahedral: Poliovirus, Helical: Tobacco mosaic virus |
| Envelope Presence | Lipid bilayer with glycoproteins, present or absent | Present: Influenza, Absent: Norovirus |
| Structural Proteins | Capsid, matrix, envelope proteins | Spike protein in Coronaviruses, Gag in Retroviruses |
| Genome Size | Ranges from ~2 kb to over 1.5 Mb in giant viruses | Hepatitis B virus (3.2 kb), Pandoravirus (>2 Mb) |
Core Concepts and Components
A virus particle, or virion, consists of three fundamental elements. First, the genome carries genetic information. Second, the capsid protects the genome and often mediates entry. Third, the envelope, when present, facilitates membrane fusion and immune evasion. The capsid is an assembly of multiple copies of one or more capsid proteins, arranged with precise symmetry. Icosahedral capsids exhibit 5 3 2 symmetry, while helical capsids form rod like structures. The envelope contains viral glycoproteins that are critical for host cell recognition. Galaxy Training Network includes workflows for analyzing viral genome sequences and predicting structural features from sequencing data.
The genome can be single stranded (ss) or double stranded (ds), positive sense or negative sense RNA, or DNA. For example, the Crimean Congo hemorrhagic fever virus has a negative sense RNA genome, and its RNA dependent RNA polymerase (RdRp) has been structurally characterized using cryo electron microscopy, revealing key catalytic domains Cell Discov. In giant viruses, such as those studied in integrative structural interactomics, the capsid architecture can be incredibly complex, incorporating hundreds of different proteins Nat Commun. This complexity challenges simple classification schemes and requires advanced computational methods.
Decision Criteria for Studying Virus Structure
When you decide how to analyze a virus structure, consider these criteria based on your research question and available data.
- Resolution needed. If your goal is to understand drug binding or catalytic mechanisms, aim for high resolution (under 3 Angstroms) using X ray crystallography or cryo EM. For assessing overall morphology or symmetry, lower resolution techniques like negative stain EM may suffice.
- Genome availability. If sequence data are available from repositories like the NCBI Sequence Read Archive, you can predict capsid protein folds using homology modeling or deep learning approaches.
- Biosafety level. Enveloped viruses with high pathogenicity, such as Crimean Congo hemorrhagic fever virus, require BSL 4 facilities for experimental work. Structural studies often rely on recombinant proteins or inactivated samples.
- Time and cost. Cryo EM is time intensive but can capture multiple conformations. X ray crystallography requires high quality crystals, which may be difficult for flexible or heterogeneous viral assemblies. Computational prediction is faster but may miss post translational modifications or dynamic states.
- Data interpretation expertise. If you lack a background in structural biology, start with curated resources from Bioconductor for genomic analysis and then move to structural databases. The framework should match your team's skill set.
Practical Workflow for Structural Analysis
Follow this workflow when you aim to determine or model a virus structure from raw data. Each step includes source bounded tools and checks.
- Sample or data acquisition. Obtain purified virus particles or express recombinant capsid proteins. For sequence based analysis, download raw reads from the NCBI Sequence Read Archive or use public genomes. Ensure sample integrity with negative stain EM or gel electrophoresis.
- Data preprocessing. For cryo EM, perform motion correction, contrast transfer function estimation, and particle picking using software like RELION. For sequences, align reads to a reference genome or assemble de novo. Quality check with fastqc.
- Structural determination. Use single particle analysis for icosahedral viruses or helical reconstruction for filamentous ones. For crystallography, solve phases with molecular replacement if a homologous structure exists. Alternatively, predict ab initio using tools from Galaxy Training Network.
- Model building and refinement. Build atomic models into electron density maps or cryo EM volumes. Refine using iterative cycles of real space refinement and validation. For RNA viruses, consider RNA structure prediction tools. A knowledge based model can assess thermodynamic stability of RNA elements J Chem Theory Comput.
- Validation and deposition. Validate using metrics like Fourier shell correlation for EM, R free for crystallography. Deposit coordinates and maps in the Protein Data Bank or EM Data Bank. Check for model completeness and ligand binding sites.
This workflow is iterative, you may need to return to data acquisition if resolution is unsatisfactory. For enveloped viruses, incorporate glycoprotein modeling from homologous structures.
Quality Checks and Validation
Rigorous validation prevents overinterpretation of virus structures. Apply these checks at each stage.
- Resolution estimates. For cryo EM, the gold standard Fourier shell correlation at 0.143 cutoff should be reported. Anything above 4 Angstroms is considered low resolution and may not permit accurate atomic modeling.
- Model geometry. Check bond lengths, angles, and Ramachandran outliers. Poor geometry indicates overfitting or map errors. The EMBL-EBI Training materials include tutorials on validation tools.
- Map model correlation. Compute cross correlation between the model and the experimental map. For crystallography, use R free and R work values. A large gap suggests model bias.
- Ligand and glycan modeling. If you model glycans on envelope proteins, ensure they fit density and are not artifacts. For drugs or inhibitors, such as imidazo pyridine derivatives against HCV NS4B, validate binding modes with independent assays ChemMedChem.
- Reproducibility. Use consistent sample preparation and data collection parameters. Compare with published structures of related viruses, such as the GIIc subtype PEDV strain BMC Vet Res, to check for consistency in capsid symmetry.
Common Mistakes in Virus Structure Interpretation
Avoid these frequent errors when working with viral assemblies.
- Assuming a single conformation. Viral proteins, especially envelope glycoproteins, adopt multiple conformational states during entry and fusion. A static structure may miss functionally relevant flexibility.
- Misinterpreting symmetry. Icosahedral viruses often have asymmetric triangulation numbers. Mistaking the T number can lead to incorrect placement of capsid subunits. Always verify with published classification.
- Overlooking the envelope. For enveloped viruses, the lipid bilayer is often not resolved in cryo EM due to disorder. Do not model the envelope unless you have clear density. Envelope proteins may also be heterogeneous due to glycosylation.
- Ignoring genome packaging. The nucleic acid inside the capsid is frequently only partially ordered. Modeling RNA or DNA based on secondary structure predictions without experimental constraints can produce misleading results.
- Using low resolution structures for drug design. If the resolution is above 3.5 Angstroms, side chain positions are uncertain. Hypotheses about drug interactions require higher resolution validation or complementary mutational data.
Limits of Uncertainty
Virus structure analysis has inherent limits that you must acknowledge.
- Resolution limits the precision of atomic coordinates. Even at 2 Angstroms, dynamic loops may be poorly resolved. Cryo EM at near atomic resolution remains rare for heterogeneous samples.
- In vitro structures may not represent in vivo states. Purification and freezing conditions can alter conformations. For example, the envelope may collapse or glycoproteins may rearrange upon extraction.
- Giant viruses challenge conventional methods. Their large genomes and complex capsids require integrative approaches that combine multiple data types, as seen in the recent analysis of a giant virus Nat Commun. Single method studies may be insufficient.
- Sequence based predictions have limited accuracy for novel folds. Homology models depend on the existence of a close template. For entirely new architectures, experimental determination is essential.
- Biological variability. Viral isolates can differ in structure due to mutations or host adaptations. The same virus strain may show heterogeneity in capsid size or envelope composition.
Frequently Asked Questions
What is the difference between a capsid and a nucleocapsid?
A capsid is the protein shell that encloses the viral genome. A nucleocapsid refers to the genome complexed with capsid proteins inside the shell. In some viruses, the nucleocapsid is a distinct substructure, especially in filamentous forms.
Can virus structure be predicted entirely from sequence?
Partial prediction is possible using homology modeling or deep learning for capsid proteins, but the quaternary assembly, symmetry, and envelope interactions require experimental validation. RNA structure can be modeled with apps like those in Bioconductor, but accuracy decreases for long genomes.
Why do some viruses have envelopes while others do not?
Envelopes provide advantages in membrane fusion and immune evasion but make viruses more fragile to environmental stresses. Non enveloped viruses often use stable capsid interactions to survive harsh conditions.
How do researchers determine the symmetry of a virus capsid?
Cryo EM single particle analysis classifies particles into symmetry groups, typically icosahedral or helical. For large complex viruses, subtomogram averaging of cryo ET data reveals local symmetry variations.
References and Further Reading
- NCBI Bookshelf for general virology textbooks and structural biology chapters.
- EMBL-EBI Training for courses on structural bioinformatics and cryo EM data analysis.
- Galaxy Training Network for workflows on viral sequence assembly and genome annotation.
- Bioconductor for R packages dedicated to genomic and structural data analysis.
- NCBI Sequence Read Archive for accessing raw sequencing data of viruses.
- Cell Discov for the structure of Crimean Congo hemorrhagic fever virus RNA dependent RNA polymerase.
- Nat Commun for integrative structural interactomics of a giant virus.
- BMC Vet Res for isolation and pathogenicity of a highly virulent recombinant PEDV strain.
- J Chem Theory Comput for thermodynamic properties of RNA molecules using a knowledge based model.
- ChemMedChem for design and binding mode analysis of imidazo pyridine derivatives against HCV.