# PDB vs. EMDB vs. AlphaFold DB: Which Structural Database Should You Use for Your Research?

Researchers in structural biology face a practical decision when choosing where to obtain macromolecular structures. The Protein Data Bank (PDB) stores experimentally determined atomic models, the Electron Microscopy Data Bank (EMDB) stores cryo-electron microscopy density maps, and the AlphaFold Database (AlphaFold DB) stores computationally predicted models. Each resource serves a different purpose, and the correct choice depends on your research question, the resolution you need, and whether you require experimental validation or broad coverage. This article provides a direct comparison of these three databases, explains their data types and limitations, and offers concrete criteria for selecting the appropriate resource or combining multiple databases in your workflow.

## At a Glance: Database Comparison Table

| Database | Data Type | Resolution or Confidence Metric | Coverage | Primary Use Case |
|----------|-----------|-------------------------------|----------|------------------|
| PDB | Experimentally determined atomic coordinates from X-ray crystallography, NMR, and cryo-EM | Resolution typically better than 4 Å for cryo-EM derived models, validation metrics include MolProbity scores | Deposited structures only, coverage depends on community submissions | Validated atomic models for docking, mutagenesis interpretation, and structure-function analysis |
| EMDB | Cryo-EM density maps, including medium-resolution maps from 5 to 10 Å | Map resolution ranges from near-atomic to low resolution, local resolution varies across the map | All deposited cryo-EM maps regardless of model quality | Visualizing conformational states, fitting models into density, and analyzing dynamic complexes |
| AlphaFold DB | Computationally predicted models using deep learning | Per-residue pLDDT confidence scores from 0 to 100, no experimental resolution | Nearly complete proteomes for many organisms | Broad sequence coverage, hypothesis generation, and template generation for difficult targets |

The table above summarizes the core distinctions. The PDB provides gold-standard experimental structures. The EMDB holds the raw density data that supports many PDB entries and also contains maps without associated atomic models. AlphaFold DB offers predicted structures for proteins that may lack experimental data entirely. Your choice among these depends on whether you prioritize experimental validation, conformational information, or sequence coverage.

## Understanding the Protein Data Bank: Experimental Gold Standard

The Protein Data Bank is the primary repository for experimentally determined three-dimensional structures of biological macromolecules. It accepts structures solved through X-ray crystallography, nuclear magnetic resonance spectroscopy, and cryo-electron microscopy. The National Center for Biotechnology Information provides access to PDB data through its structure search systems, which integrate sequence and structure information for researchers who need to locate specific entries or compare related proteins.

### Data Types and Quality Metrics in the PDB

PDB entries contain atomic coordinates, experimental details, and validation metrics. For cryo-EM derived models, the resolution of the underlying density map is a critical quality indicator. Models solved from experimental densities with resolution higher than 4 Å are generally considered reliable for detailed structural interpretation. When you download a PDB entry, you should examine the resolution, R-free values for crystallographic structures, and MolProbity scores that assess geometric quality including clash scores and Ramachandran outliers.

### When to Use the PDB

Use the PDB when you need experimentally validated atomic coordinates for molecular docking, site-directed mutagenesis planning, or structure-based drug design. The PDB is also the appropriate choice when you need to compare your own experimental structure against known structures of related proteins. For example, researchers studying the HIV-1 envelope glycoprotein rely on PDB entries from cryo-EM studies to understand conformational states that are relevant for vaccine design. These experimental structures capture the atomic details of the envelope protein in closed and open states, providing a structural roadmap for targeting dynamic states.

### Limitations of the PDB

The PDB has a fundamental coverage limitation. It only contains structures that researchers have experimentally determined and deposited. Many proteins, particularly membrane proteins, large complexes, and proteins from less-studied organisms, remain absent from the PDB. Additionally, some PDB entries are derived from medium-resolution cryo-EM maps where the atomic model may contain inaccuracies. A systematic study comparing PDB structures with AlphaFold predictions found that structural segments derived from medium-resolution cryo-EM maps often have poorer MolProbity scores than corresponding AlphaFold predictions, suggesting that some PDB models could benefit from refinement using predicted structures.

## Understanding the EMDB: Density Maps for Conformational Analysis

The Electron Microscopy Data Bank stores three-dimensional density maps determined by cryo-electron microscopy. Unlike the PDB, which stores atomic models, the EMDB stores the actual density data that shows where atoms are located within a macromolecular complex. This distinction matters because density maps contain information that atomic models may not fully capture.

### Data Types in the EMDB

EMDB entries include the three-dimensional density map, information about the microscope and detector used, the number of particles averaged, and the final resolution estimate. Maps can range from near-atomic resolution below 3 Å to low-resolution maps at 10 Å or worse. Medium-resolution maps between 5 and 10 Å are increasingly common as cryo-EM becomes more accessible to the broader research community.

### When to Use the EMDB

Use the EMDB when you need to examine conformational heterogeneity, visualize flexible regions that may not be well-defined in atomic models, or fit existing atomic models into experimental density. The EMDB is also essential when you want to understand the quality of the experimental data that supports a PDB entry. For example, researchers studying the fission yeast phosphate exporter SpXpr1 used cryo-EM to determine structures in both the apo and inositol hexaphosphate bound states. The density maps in the EMDB allow other researchers to independently assess the structural interpretations and examine regions that may be ambiguous in the atomic model.

### Limitations of the EMDB

The EMDB contains maps of varying quality, and not all maps have associated atomic models. Medium-resolution maps between 5 and 10 Å are difficult to interpret directly because individual amino acid side chains are not visible. Researchers often fit known templates or predicted structures into these maps, which introduces model bias. The EMDB does not provide validation scores for the biological interpretation of the map, only technical metrics about map quality.

## Understanding AlphaFold DB: Predicted Models for Broad Coverage

The AlphaFold Protein Structure Database contains computationally predicted protein structures generated by the AlphaFold deep learning system. The database provides predicted structures for nearly complete proteomes of many organisms, offering coverage that far exceeds the experimental PDB.

### Data Types and Confidence Metrics in AlphaFold DB

AlphaFold DB entries contain predicted atomic coordinates along with per-residue confidence scores. The predicted local distance difference test (pLDDT) score ranges from 0 to 100 and indicates the model confidence for each residue. High-confidence regions with pLDDT above 90 are suitable for detailed analysis, while low-confidence regions below 50 should be treated with caution. The database also provides predicted aligned error (PAE) plots that indicate the relative positional confidence between residue pairs, which is useful for assessing domain arrangements.

### When to Use AlphaFold DB

Use AlphaFold DB when you need structures for proteins that lack experimental data, when you want to generate hypotheses about protein function, or when you need templates for molecular replacement in crystallography. AlphaFold predictions have proven valuable for refining models derived from medium-resolution cryo-EM maps. A study that systematically mapped paired structural segments between the PDB and AlphaFold DB found that AlphaFold predictions had better MolProbity scores on average than the corresponding PDB segments derived from medium-resolution cryo-EM data. This suggests that AlphaFold models can serve as improved starting points for interpreting medium-resolution density maps.

### Limitations of AlphaFold DB

AlphaFold predictions are computational models, not experimental measurements. They do not capture ligand binding, post-translational modifications, or conformational changes induced by environmental conditions. The models represent a predicted ground state and may miss alternative conformations that are biologically important. For example, the HIV-1 envelope glycoprotein undergoes dramatic conformational changes during host cell entry, and these dynamic states cannot be captured by a single prediction. AlphaFold models also have lower confidence in regions that lack evolutionary constraints, such as flexible loops and disordered regions.

## Comparing Data Quality and Reliability Across Databases

The reliability of structural data varies substantially across the three databases. Experimental structures in the PDB are generally considered the gold standard because they are validated against actual measurements. However, the quality of PDB entries varies, particularly for structures derived from medium-resolution cryo-EM maps.

### Resolution and Confidence Metrics

The PDB provides resolution information that directly indicates the level of detail visible in the experimental data. Cryo-EM structures at resolutions better than 4 Å allow confident placement of amino acid side chains. Medium-resolution maps between 5 and 10 Å only show the overall protein fold and secondary structure elements. AlphaFold DB does not provide a resolution value because the models are predictions. Instead, the pLDDT score serves as the confidence metric, with values above 90 indicating high confidence and values below 50 indicating low confidence.

### Validation Scores and Their Interpretation

MolProbity is a widely used validation tool that assesses the geometric quality of atomic models. It evaluates clash scores, Ramachandran outliers, and rotamer outliers. A systematic comparison of PDB and AlphaFold DB structural segments derived from medium-resolution cryo-EM maps found that AlphaFold segments had an average MolProbity score of 0.96, while the corresponding PDB segments averaged 1.98. Lower MolProbity scores indicate better geometry. This finding suggests that AlphaFold predictions often have better local geometry than models manually built into medium-resolution density maps.

### Coverage Comparison

The PDB contains approximately 200,000 experimentally determined structures, while AlphaFold DB contains hundreds of millions of predicted structures covering entire proteomes. The EMDB contains over 30,000 cryo-EM maps. This dramatic difference in coverage means that AlphaFold DB is often the only source of structural information for less-studied proteins. However, coverage alone does not determine which database you should use. The research question determines the appropriate resource.

## Practical Workflow: Selecting the Right Database for Your Research Question

The following workflow provides concrete steps for deciding which database to use based on your specific research needs.

### Step 1: Define Your Research Question

Write down exactly what you need to know about your protein of interest. Are you planning site-directed mutagenesis experiments? Are you interpreting the results of a docking study? Are you analyzing conformational changes during a biological process? Are you studying a protein that has no experimental structure? The answers to these questions determine which database is most appropriate.

### Step 2: Check the PDB First

Search the PDB for your protein of interest. If a high-resolution experimental structure exists, use it as your primary resource. High-resolution structures from X-ray crystallography or cryo-EM at resolutions better than 3 Å provide reliable atomic details for most applications. Check the resolution and validation metrics for any PDB entry you plan to use.

### Step 3: Examine the EMDB for Conformational Information

If your research involves conformational dynamics, ligand-induced changes, or complex assembly, search the EMDB for density maps of your protein. Even if a PDB entry exists, the EMDB may contain maps of different conformational states that provide additional biological insight. For example, the HIV-1 envelope glycoprotein has been captured in multiple conformational states using cryo-EM, and these maps reveal the stepwise rearrangements that occur during host cell entry.

### Step 4: Use AlphaFold DB for Missing Coverage

If no experimental structure exists for your protein, or if you need structures for many related proteins, use AlphaFold DB. Check the pLDDT scores for the regions you plan to analyze. High-confidence regions can support detailed analysis, while low-confidence regions require experimental validation before drawing conclusions.

### Step 5: Combine Databases When Appropriate

For many research questions, the best approach combines multiple databases. You might use an AlphaFold prediction as a starting model for fitting into a medium-resolution cryo-EM map from the EMDB. You might compare an AlphaFold prediction against a PDB structure to identify conformational differences. You might use the EMDB to examine density features that are not well-modeled in the PDB entry.

## Options and Tradeoffs: Choosing Between Experimental and Predicted Structures

Each database offers distinct advantages and disadvantages that affect your research outcomes.

### Experimental Structures: The Tradeoff Between Quality and Coverage

Experimental structures in the PDB provide validated atomic coordinates that have been checked against actual measurements. These structures are essential for understanding precise atomic interactions, planning mutagenesis experiments, and interpreting functional data. The tradeoff is limited coverage. Many proteins remain unsolved because they are difficult to express, purify, or crystallize. Membrane proteins and large dynamic complexes are particularly underrepresented in the PDB.

### Density Maps: The Tradeoff Between Detail and Interpretability

EMDB maps provide the raw experimental data from cryo-EM experiments. High-resolution maps below 3 Å allow de novo model building with confidence. Medium-resolution maps between 5 and 10 Å show the overall fold but not atomic details. These maps require fitting of known structures or predicted models, which introduces potential bias. The tradeoff is that density maps capture conformational states and heterogeneity that atomic models may miss, but interpreting them requires additional modeling steps.

### Predicted Models: The Tradeoff Between Coverage and Validation

AlphaFold DB provides predicted structures for nearly entire proteomes, offering coverage that experimental methods cannot match. These predictions are valuable for hypothesis generation, identifying functional residues, and planning experiments. The tradeoff is that predictions lack experimental validation. They do not capture ligand-induced conformational changes, post-translational modifications, or the effects of mutations. Predictions are also less reliable for regions with low evolutionary conservation, such as flexible loops and disordered regions.

## Observations and Measurements: What to Record in Your Structural Analysis

When you use structural databases in your research, maintain detailed records of your data sources and analysis steps. This documentation supports reproducibility and allows you to trace conclusions back to specific data.

### Record the Database Version and Entry Identifiers

Record the exact PDB ID, EMDB ID, or AlphaFold DB identifier for every structure you use. Note the database release date and version, because structures may be updated or superseded. For AlphaFold DB, record the model organism and the specific isoform or splice variant you used.

### Record Quality Metrics for Each Structure

For PDB entries, record the resolution, R-free value, and MolProbity score. For EMDB entries, record the map resolution and the number of particles used in the reconstruction. For AlphaFold DB entries, record the pLDDT scores for the regions you analyze and note any low-confidence regions.

### Record Your Analysis Parameters

Document the software and parameters you used for structure alignment, docking, or model fitting. Record the version numbers of all software packages. This information allows other researchers to reproduce your analysis and assess the reliability of your conclusions.

## Quality Controls and Validation Checks

Implement quality controls at each stage of your structural analysis to avoid drawing incorrect conclusions from unreliable data.

### Check Resolution Before Detailed Interpretation

Do not interpret atomic details from structures with resolution worse than 4 Å. At medium resolution between 5 and 10 Å, only the overall fold and secondary structure are reliable. Side chain positions and specific atomic interactions should not be interpreted from such data.

### Check Confidence Scores in AlphaFold Predictions

Before using an AlphaFold prediction for detailed analysis, examine the pLDDT scores across the protein. Regions with pLDDT below 70 should be treated with caution. Regions below 50 are unreliable and should not be used for structural interpretation without experimental validation.

### Cross-Validate Predictions Against Experimental Data

When possible, compare AlphaFold predictions against any available experimental data. If a PDB structure exists for a homologous protein, compare the predicted structure against the experimental structure to assess the reliability of the prediction. If you have biochemical data about residue function, check whether functionally important residues are in high-confidence regions of the prediction.

### Validate Model Fitting into Density Maps

When fitting models into medium-resolution cryo-EM maps, use multiple independent fitting approaches and compare the results. Check that the fitted model is consistent with known biochemical constraints, such as disulfide bonds, glycosylation sites, and protein-protein interaction interfaces.

## Common Failure Patterns in Structural Database Use

Researchers commonly encounter several recurring problems when using structural databases. Recognizing these patterns helps you avoid them.

### Overinterpreting Low-Resolution Structures

A frequent error is treating medium-resolution cryo-EM structures as if they were high-resolution structures. Researchers may interpret side chain conformations or specific hydrogen bonds from structures at 6 or 7 Å resolution, where such details are not visible in the density. This leads to incorrect conclusions about molecular mechanisms.

### Ignoring AlphaFold Confidence Scores

Another common failure is using AlphaFold predictions without checking pLDDT scores. Low-confidence regions may have incorrect fold predictions, and using these regions for detailed analysis produces unreliable results. Always check confidence scores before drawing conclusions from predicted structures.

### Assuming PDB Structures Are Always Correct

PDB entries are experimental results, but they can contain errors. Models built into medium-resolution maps may have incorrect register shifts or misplaced loops. Always check validation metrics and examine the density map in the EMDB when available.

### Neglecting Conformational Diversity

A single structure, whether experimental or predicted, represents one conformational state. Proteins are dynamic molecules that adopt multiple conformations. Using a single structure to interpret all functional data can lead to incorrect conclusions about mechanisms that involve conformational changes.

## Limitations and Interpretation Boundaries

Every structural database has limitations that affect how you can interpret the data.

### PDB Limitations

The PDB contains structures determined under specific experimental conditions that may not reflect the physiological environment. Crystal packing forces, detergent micelles, or grid preparation artifacts can influence the observed conformation. The PDB also has limited representation of certain protein classes, particularly membrane proteins and intrinsically disordered proteins.

### EMDB Limitations

EMDB maps represent averaged reconstructions from many particles. Flexible regions may be poorly resolved because they adopt multiple conformations in the sample. The map resolution is a global estimate, and local resolution can vary substantially across the map. Some regions may be well-resolved while others are poorly defined.

### AlphaFold DB Limitations

AlphaFold predictions are based on evolutionary information from multiple sequence alignments. Proteins with few homologs may have lower prediction accuracy. The predictions do not account for environmental conditions, binding partners, or post-translational modifications. The models represent a single predicted conformation and do not capture the dynamic behavior of proteins.

## Safety and Reproducibility Context

Structural biology research increasingly relies on computational resources and reproducible workflows. Several training resources provide guidance for developing reproducible analysis pipelines.

### Reproducible Workflow Training

The Galaxy Training Network offers accessible tutorials for bioinformatics analysis, including structural biology workflows. These tutorials emphasize reproducible analysis through documented workflows that can be shared and rerun. The nf-core documentation describes community standards for building reproducible analysis pipelines using Nextflow. These resources help researchers implement quality controls and documentation practices in their structural analysis workflows.

### Foundational Computing Skills

The Carpentries provides lessons on foundational computing skills including shell, Git, and programming. These skills are essential for managing structural data files, versioning analysis scripts, and documenting computational workflows. The EMBL-EBI Training portal offers learning pathways specifically for bioinformatics data resources, including training on structural databases and their use.

### Bioconductor for Structural Analysis

Bioconductor provides R packages for genomic and structural analysis. These packages support reproducible analysis through documented workflows and versioned software. Researchers can use Bioconductor tools to integrate structural data with other biological data types.

## Professional Escalation Criteria

Some situations require consultation with structural biology experts or specialized support services.

### When to Seek Expert Help

Consult a structural biology expert when you need to interpret medium-resolution cryo-EM maps, when you are planning to build atomic models into density maps, or when you need to assess the reliability of a specific structure for your research. Experts can help you evaluate map quality, assess model bias, and determine whether your planned analysis is appropriate for the available data.

### When to Use Specialized Training

If you are new to structural biology, complete training through the EMBL-EBI Training portal or the Galaxy Training Network before conducting detailed structural analysis. These resources provide structured learning pathways that cover the fundamentals of structural data interpretation.

### When to Contact Database Curators

Contact database curators when you identify potential errors in database entries, when you need assistance accessing specific data, or when you have questions about data formats. The NCBI provides support for its structure resources, and the PDB and EMDB have dedicated help desks.

## A Decision Framework for Matching Structural Databases to Research Questions

The preceding sections described what each database contains and how to assess data quality. This section provides a structured decision framework that maps specific research questions to the appropriate database or combination of databases. The framework uses a scoring system that weighs experimental validation, conformational coverage, and sequence coverage against the demands of your research question.

### The Three-Axis Scoring Method

Define your research question along three axes before selecting a database. Each axis receives a score from 1 to 5 based on how strongly your research depends on that type of information.

**Axis 1: Experimental Validation Demand**

Score this axis based on how critically your conclusions depend on experimentally verified atomic positions. Score 5 when you need precise atomic coordinates for drug design, mutagenesis planning, or mechanism interpretation. Score 3 when you need a general structural context but will validate specific interactions through biochemical experiments. Score 1 when you only need a rough structural framework for hypothesis generation.

**Axis 2: Conformational Diversity Requirement**

Score this axis based on whether your research question involves dynamic processes, multiple functional states, or ligand-induced changes. Score 5 when you study conformational transitions, allosteric regulation, or assembly mechanisms. Score 3 when you need a representative structure but acknowledge the protein may adopt multiple states. Score 1 when a single static structure adequately represents the state you study.

**Axis 3: Sequence Coverage Necessity**

Score this axis based on how many proteins or homologs you need to analyze. Score 5 when you study entire families, compare many orthologs, or need structures for proteins without experimental data. Score 3 when you study a small number of related proteins. Score 1 when you focus on a single well-characterized protein.

After scoring each axis, use the following decision rules to select your primary database.

### Decision Rule 1: High Experimental Validation Demand

When your experimental validation demand scores 4 or 5, the PDB serves as your primary database. This applies to structure-based drug design, interpretation of mutagenesis data, and analysis of precise atomic interactions. Search the PDB first and examine the resolution and validation metrics for candidate entries. If a high-resolution structure exists at better than 3 Å, use it directly. If only medium-resolution structures exist, note their limitations and consider whether the EMDB map provides additional information for interpreting ambiguous regions.

For example, researchers studying the glycine receptor mGlyR for depression therapy solved an atomic structure of the receptor bound to a nanobody using cryo-EM. This experimental structure provides the precise atomic details needed to understand the mechanism of receptor modulation and to guide further biologic development. A predicted model would not capture the binding interface between the nanobody and receptor with sufficient accuracy for this purpose.

### Decision Rule 2: High Conformational Diversity Requirement

When your conformational diversity requirement scores 4 or 5, the EMDB becomes essential even if a PDB entry exists. The EMDB contains density maps of multiple conformational states that atomic models may not fully capture. This applies to studies of membrane transporters, signaling complexes, viral glycoproteins, and other dynamic systems.

The HIV-1 envelope glycoprotein provides a clear example. Cryo-EM and cryo-electron tomography studies have resolved the envelope protein in closed and multiple open states, revealing the stepwise structural rearrangements that occur during host cell entry. Researchers established a classification framework for symmetric and asymmetric envelope states based on their degree of openness. A single PDB entry or AlphaFold prediction cannot represent this conformational landscape. The EMDB maps provide the raw density data that supports analysis of these distinct states.

Similarly, the fission yeast phosphate exporter SpXpr1 was studied using cryo-EM structures in both the apo and inositol hexaphosphate bound states. The density maps in the EMDB allow researchers to examine the dual gating mechanism involving the intracellular N-loop gate and the extracellular plug. These conformational states are only accessible through experimental density maps.

### Decision Rule 3: High Sequence Coverage Necessity

When your sequence coverage necessity scores 4 or 5, AlphaFold DB provides the primary resource. This applies to comparative studies across species, analysis of protein families, and research on proteins that lack experimental structures. AlphaFold DB offers predicted structures for nearly complete proteomes, enabling analyses that would be impossible with experimental data alone.

Check the pLDDT scores for the regions you plan to analyze. High-confidence regions above 90 support detailed structural comparison. Regions between 70 and 90 support fold-level analysis but require caution for atomic details. Regions below 50 should be excluded from structural interpretation.

### Decision Rule 4: Balanced Scores Require Database Combination

When your scores are balanced across axes, combine databases in a documented workflow. The most common combination uses AlphaFold predictions as starting models for fitting into medium-resolution cryo-EM maps from the EMDB. This approach addresses the difficulty of building atomic models directly from maps at 5 to 10 Å resolution.

A systematic study of paired structural segments between the PDB and AlphaFold DB examined 918 nonredundant pairs derived from medium-resolution cryo-EM maps. The AlphaFold segments achieved an average MolProbity score of 0.96, while the corresponding PDB segments averaged 1.98. Lower MolProbity scores indicate better geometry. This finding supports using AlphaFold predictions as improved starting points for interpreting medium-resolution density maps, particularly when the existing PDB model has poor geometric quality.

### Implementing the Framework in Your Research Workflow

Apply the framework systematically by documenting your axis scores and database selection before conducting structural analysis. This documentation supports reproducibility and provides a clear rationale for your choices.

**Step 1: Score Your Research Question**

Write your research question and assign scores for experimental validation demand, conformational diversity requirement, and sequence coverage necessity. Record these scores in your laboratory notebook or electronic lab notebook.

**Step 2: Select Your Primary Database**

Apply the decision rules to select your primary database. Record the database name and the rationale for your selection.

**Step 3: Identify Candidate Entries**

Search your primary database for relevant entries. Record the entry identifiers, resolution or confidence metrics, and any relevant validation scores.

**Step 4: Assess Whether Supplementary Databases Are Needed**

Evaluate whether your primary database fully addresses your research question. If conformational diversity is important but your primary database is the PDB, search the EMDB for additional maps. If experimental validation is important but your primary database is AlphaFold DB, search the PDB for any experimental structures of your protein or close homologs.

**Step 5: Document Your Final Database Set**

Record the complete set of database entries used in your analysis, including the role each entry plays. This documentation supports reproducibility and allows reviewers to assess the appropriateness of your data sources.

### A Worked Example: Studying a Membrane Transporter

Consider a researcher studying a membrane transporter involved in phosphate homeostasis. The research question asks how the transporter regulates phosphate export and how inositol hexaphosphate binding affects its function.

**Axis scoring:** Experimental validation demand scores 4 because the researcher needs atomic details of the transport mechanism. Conformational diversity requirement scores 5 because the transporter likely adopts different states during the transport cycle. Sequence coverage necessity scores 2 because the researcher focuses on one transporter but may compare with orthologs.

**Database selection:** The high conformational diversity score directs the researcher to the EMDB as the primary database. The researcher searches for cryo-EM maps of the transporter and identifies entries for both the apo and ligand-bound states. The PDB serves as a supplementary database for atomic models derived from these maps. AlphaFold DB provides a useful comparison for assessing whether the predicted structure matches the experimentally determined states.

**Analysis approach:** The researcher examines the EMDB maps to assess the density quality in the gate regions. The PDB models provide atomic coordinates for interpreting specific interactions. The AlphaFold prediction offers a reference for evaluating whether the experimental structures reveal conformational changes not captured by prediction.

### Common Errors in Database Selection

Several recurring errors undermine structural analysis. Recognizing these patterns helps you apply the framework correctly.

**Error 1: Using AlphaFold Predictions for Conformational Studies**

AlphaFold predictions represent a single predicted conformation and do not capture the dynamic behavior of proteins. Researchers studying conformational transitions sometimes use AlphaFold models as if they represented specific functional states. This leads to incorrect conclusions about mechanisms that involve structural rearrangements. The HIV-1 envelope glycoprotein example illustrates this problem. The envelope protein transitions from closed to fully open states during host cell entry, and these states cannot be captured by a single prediction.

**Error 2: Ignoring the EMDB When PDB Entries Exist**

Researchers often use PDB atomic models without examining the underlying density maps in the EMDB. This practice can hide model errors, particularly in regions where the density is weak or ambiguous. The EMDB map provides the raw experimental data that allows independent assessment of the model interpretation. For medium-resolution maps between 5 and 10 Å, examining the density is essential for understanding which regions of the model are reliable.

**Error 3: Treating All PDB Entries as Equal Quality**

PDB entries vary substantially in quality. Structures derived from medium-resolution cryo-EM maps may have poorer geometry than structures solved by high-resolution crystallography. The systematic comparison of PDB and AlphaFold DB segments found that PDB segments from medium-resolution cryo-EM maps had an average MolProbity score of 1.98, indicating notable geometric issues. Always check the resolution and validation metrics before using a PDB entry for detailed analysis.

**Error 4: Failing to Document Database Versions**

Structural databases are updated regularly. Entries may be superseded, corrected, or removed. Failing to record the database version and entry identifiers makes it difficult to reproduce your analysis or trace conclusions back to specific data. Record the exact identifiers and access dates for all database entries used in your research.

### Records and Measurements for Database Selection

Maintain a structured record of your database selection process and the data you use. This record supports reproducibility and provides evidence for the appropriateness of your choices.

**Database Selection Record**

Record your axis scores, the decision rules applied, and the resulting database selection. Include the date of your analysis and the version of any software tools used.

**Entry Quality Record**

For each database entry used, record the identifier, resolution or confidence metrics, validation scores, and any relevant notes about data quality. For PDB entries, record the resolution, R-free value, and MolProbity score. For EMDB entries, record the map resolution and particle number. For AlphaFold DB entries, record the pLDDT scores for the regions analyzed.

**Analysis Parameter Record**

Document the software and parameters used for structure alignment, model fitting, or other analyses. Record version numbers for all software packages. This information allows other researchers to reproduce your analysis.

### Troubleshooting Database Selection Problems

When your structural analysis produces unexpected results, review your database selection using a structured troubleshooting approach.

**Problem 1: Predicted and Experimental Structures Disagree**

When an AlphaFold prediction differs substantially from an experimental structure, examine the confidence scores in the prediction and the resolution of the experimental structure. Low pLDDT scores in the divergent regions indicate the prediction is unreliable in those areas. Poor resolution in the experimental structure may indicate model errors. Compare the structures to identify which regions diverge and assess the evidence for each conformation.

**Problem 2: Density Map Quality Is Insufficient for Interpretation**

When a cryo-EM map does not support the level of interpretation you need, check the global and local resolution estimates. Medium-resolution maps between 5 and 10 Å only support fold-level interpretation. If you need atomic details, search for higher-resolution structures of your protein or closely related homologs. Consider whether the EMDB contains maps of different conformational states that may have better local resolution in the regions you study.

**Problem 3: No Structure Exists for Your Protein**

When your protein is absent from all three databases, search for structures of homologous proteins in the PDB. Use the NCBI structure search systems to identify related entries. Generate an AlphaFold prediction using the AlphaFold software if no prediction exists in the database. Assess the confidence of the prediction and use it as a hypothesis-generating tool instead of a definitive structure.

**Problem 4: Validation Scores Indicate Poor Model Quality**

When a PDB entry has poor MolProbity scores or other validation metrics, examine the EMDB map to assess whether the model fits the density. Consider whether an AlphaFold prediction provides a better starting model for the regions with poor geometry. The systematic comparison of PDB and AlphaFold DB segments suggests that AlphaFold predictions often have better geometry than models built into medium-resolution maps.

### Escalation Criteria for Database Selection Decisions

Some situations require consultation with structural biology experts or specialized training before proceeding with analysis.

**Escalate When You Need to Build Models into Medium-Resolution Maps**

Building atomic models into maps at 5 to 10 Å resolution requires specialized expertise. The fitting process introduces model bias, and incorrect interpretations can propagate through downstream analysis. Consult a structural biologist who has experience with medium-resolution map interpretation before attempting this analysis.

**Escalate When You Plan Structure-Based Drug Design**

Structure-based drug design requires high-quality experimental structures with accurate atomic positions. If only medium-resolution structures or predicted models are available, consult with computational chemists or structural biologists about the limitations and whether additional experimental work is needed.

**Escalate When You Encounter Conflicting Structural Data**

When different database entries or different experimental structures conflict, consult an expert to assess which structure is more reliable. Consider the resolution, validation metrics, and experimental conditions for each entry. The expert can help you evaluate the evidence and determine the appropriate interpretation.

### Training Resources for Database Selection Skills

Developing proficiency in structural database selection requires training in both structural biology concepts and computational skills. Several official training resources provide structured learning pathways.

The EMBL-EBI Training portal offers learning pathways specifically for bioinformatics data resources, including training on structural databases and their appropriate use. These courses cover the fundamentals of structural data interpretation and database selection.

The Galaxy Training Network provides accessible tutorials for bioinformatics analysis, including structural biology workflows. These tutorials emphasize reproducible analysis through documented workflows that can be shared and rerun. The nf-core documentation describes community standards for building reproducible analysis pipelines using Nextflow.

The Carpentries provides lessons on foundational computing skills including shell, Git, and programming. These skills are essential for managing structural data files, versioning analysis scripts, and documenting computational workflows. Bioconductor provides R packages for genomic and structural analysis that support reproducible analysis through documented workflows and versioned software.

### Applying the Framework to Emerging Research Areas

The decision framework applies to emerging research areas where structural biology plays an increasingly important role.

**Vaccine Design and Immunogen Engineering**

Researchers developing vaccines against SARS-CoV-2 and other viruses use structural information to design immunogens that focus antibody responses on specific epitopes. The immune-focused RBD nanoparticle approach used structure-guided design to focus antibody responses to the receptor binding site. This work required experimental structures of the spike protein in various conformations, density maps showing the epitope architecture, and predicted models for designing the nanoparticle components. The framework directs researchers to the EMDB for conformational information, the PDB for validated atomic models, and AlphaFold DB for designing engineered components.

**Biologic Development for Neuropsychiatric Conditions**

The development of nanobodies for depression treatment required solving the atomic structure of the glycine receptor bound to the nanobody. This experimental structure provided the mechanistic understanding needed to explain how the nanobody modulates receptor function. The framework directs researchers to the PDB for the experimentally validated complex structure and to the EMDB for the density maps supporting the model.

**Evolutionary Studies of Protein Families**

Researchers studying the evolution of phosphate transporters across eukaryotes used cryo-EM structures of the fission yeast transporter to understand conserved and unique features. The framework directs researchers to AlphaFold DB for comparing structures across many species, to the PDB for experimentally validated structures of key family members, and to the EMDB for examining conformational states.

### Limitations of the Decision Framework

The decision framework provides guidance but does not replace scientific judgment. Several limitations should be acknowledged.

The axis scoring system is subjective and requires researchers to assess their own needs. Different researchers may score the same research question differently. The framework works best when applied with careful consideration of the specific research context.

The framework assumes that appropriate database entries exist. When no entries are available in any database, the framework cannot guide selection. In these cases, researchers must consider experimental structure determination or comparative modeling approaches.

The framework does not address all possible research questions. Some questions may require specialized approaches not covered by the three axes. Researchers should adapt the framework to their specific needs and document any modifications.

### Integrating the Framework with Reproducible Workflow Practices

The decision framework supports reproducible research by making database selection explicit and documented. Integrate the framework with reproducible workflow practices to ensure your structural analysis can be verified and extended by other researchers.

Document your axis scores and database selection in your analysis workflow. Use version control to track changes to your analysis scripts and documentation. Share your workflow through platforms that support reproducible analysis, such as Galaxy or nf-core pipelines.

The NCBI provides access to structure data through its search systems, which integrate sequence and structure information. These systems support reproducible access to structural data by providing stable identifiers and versioned entries.

### Professional Judgment in Database Selection

The decision framework provides structure, but professional judgment remains essential. Experienced structural biologists develop intuition about data quality and appropriate use that goes beyond formal scoring systems. Use the framework as a starting point and refine your approach based on experience and feedback from colleagues.

When in doubt about database selection, consult colleagues with structural biology expertise. Discuss your research question and the available data. These discussions often reveal considerations that formal frameworks miss, such as the biological context of the structure, the specific experimental conditions, or the known limitations of particular entries.

### Summary of the Decision Framework

The three-axis scoring method provides a systematic approach to database selection. Score your research question on experimental validation demand, conformational diversity requirement, and sequence coverage necessity. Apply the decision rules to select your primary database. Combine databases when your scores are balanced across axes. Document your selection process and the quality metrics of all entries used. Escalate to expert consultation when you encounter medium-resolution map interpretation, conflicting structural data, or structure-based drug design with limited experimental data.

## Frequently Asked Questions

### What is the main difference between PDB and EMDB?

The PDB stores atomic models with coordinates for each atom in the structure. The EMDB stores the three-dimensional density maps from cryo-EM experiments. A PDB entry derived from cryo-EM has an associated EMDB map that shows the actual experimental density. The PDB model is an interpretation of the density, while the EMDB map is the raw experimental data.

### Can I use AlphaFold predictions instead of experimental structures?

AlphaFold predictions are useful when no experimental structure exists, but they are not substitutes for experimental validation. Predictions lack information about ligand binding, post-translational modifications, and conformational changes. Use AlphaFold predictions for hypothesis generation and planning experiments, but confirm critical structural details with experimental methods.

### How do I know if an AlphaFold prediction is reliable?

Check the per-residue confidence scores (pLDDT) in the AlphaFold DB entry. Regions with scores above 90 are highly reliable, scores between 70 and 90 are generally reliable, and scores below 50 are unreliable. Also examine the predicted aligned error plot to assess the confidence in domain arrangements.

### What resolution do I need for detailed structural interpretation?

For confident placement of amino acid side chains and interpretation of specific atomic interactions, use structures with resolution better than 3 Å. At resolutions between 3 and 4 Å, the backbone is reliable but side chain positions may be uncertain. At resolutions worse than 4 Å, only the overall fold and secondary structure should be interpreted.

### Why are some PDB structures derived from medium-resolution cryo-EM maps of lower quality?

Models built into medium-resolution maps between 5 and 10 Å require fitting of known templates or predicted structures because individual amino acids are not visible in the density. This fitting process can introduce errors, particularly in loop regions and areas with weak density. Validation scores such as MolProbity can identify models with poor geometry.

### How can I combine AlphaFold predictions with cryo-EM density maps?

AlphaFold predictions can serve as starting models for fitting into medium-resolution cryo-EM maps. The predicted structure provides a complete atomic model that can be flexibly fitted into the density. This approach often produces better models than fitting existing PDB structures, particularly when the PDB structure is from a different conformation or has poor geometry.

### What should I do if my protein of interest is not in any structural database?

If your protein has no experimental structure and no AlphaFold prediction, you can generate a prediction using the AlphaFold software yourself. You can also search for structures of homologous proteins in the PDB and use comparative modeling approaches. If your protein is biologically important, consider determining its structure experimentally using cryo-EM or X-ray crystallography.

### How do I cite structural database entries in my publications?

Cite the original publication that describes the structure determination, beyond the database entry. Each PDB entry includes a primary citation. For AlphaFold DB, cite the AlphaFold paper and include the specific database version and entry identifier. Check the citation guidelines provided by each database for the required format.

## Related Bioinformatics Guides

- [Digital Pathology Scanners: A Buyer's Guide for Clinical and Research Use](/knowledge/bioinformatics/digital-pathology-scanners-a-buyer-s-guide-for-clinical-and-research-use)
- [Persistent Identifiers for Research Data: A Guide to Selection and Use](/knowledge/bioinformatics/persistent-identifiers-for-research-data-a-guide-to-selection-and-use)
- [Spatial Proteomics Method of the Year: What It Means for Your Research](/knowledge/bioinformatics/spatial-proteomics-method-of-the-year-what-it-means-for-your-research)
- [Spatial Transcriptomics vs. Single-Cell RNA Sequencing: Which Approach Fits Your Research?](/knowledge/bioinformatics/spatial-transcriptomics-vs-single-cell-rna-sequencing-which-approach-fits-your-research)
- [Detecting Structural Variants with Long-Read Sequencing: Methods and Considerations](/knowledge/bioinformatics/detecting-structural-variants-with-long-read-sequencing-methods-and-considerations)

## References and Further Reading

- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Bioconductor](https://bioconductor.org/). Bioconductor Project.
- [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.
- [nf-core Documentation](https://nf-co.re/docs). nf-core.
- [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.
- [Structural insights into the gating mechanism of the fission yeast phosphate exporter SpXpr1.](https://doi.org/10.1038/s41421-026-00883-8). 2026.
- [A Data Set of Paired Structural Segments Between Protein Data Bank and AlphaFold DB for Medium-Resolution Cryo-EM Density Maps: A Gap in Overall Structural Quality.](https://doi.org/10.1007/978-981-97-5087-0_5). 2024.
- [Conformational landscape of HIV-1 Env from closed to fully open.](https://doi.org/10.1038/s41467-026-69921-z). 2026.
- [Immune-focused RBD nanoparticles induce cross-reactive, RBS-directed responses capable of variant-resistant SARS-CoV-2 neutralization.](https://doi.org/10.1371/journal.ppat.1013905). 2026.
- [Targeting mGlyR with nanobodies for depression.](https://doi.org/10.1038/s41467-026-68339-x). 2026.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.