# MolProbity vs. QMEAN: Choosing the Right Tool for Validating X-ray and Predicted Structures

Researchers working with protein structures face a fundamental decision: which validation tool to apply when assessing model quality. The choice depends primarily on the origin of the structure. Experimental structures determined by X-ray crystallography require validation against geometric and stereochemical criteria that reflect the physical constraints of protein folding. Predicted structures generated by homology modeling or deep learning methods require evaluation against statistical potentials derived from known protein structures. MolProbity and QMEAN represent these two distinct validation philosophies. MolProbity evaluates experimental structures through all-atom contact analysis and geometric criteria. QMEAN assesses predicted models through composite scoring against reference structures. This article explains the underlying principles, input requirements, score interpretation, and practical decision criteria for both tools, with specific guidance for researchers working across the structural biology workflow.

## The Structural Validation Problem

Protein structure determination produces a model that must be assessed for reliability before it can support biological conclusions. A structure file, whether deposited in the Protein Data Bank or generated by prediction software, contains atomic coordinates that describe a proposed three-dimensional arrangement of amino acid residues. These coordinates carry inherent uncertainty. Experimental structures contain errors from diffraction data quality, refinement procedures, and model building decisions. Predicted structures contain errors from template selection, alignment accuracy, and the inherent limitations of prediction algorithms.

Validation tools address this uncertainty by applying mathematical and statistical tests to the coordinate data. The tests evaluate whether the proposed structure conforms to known physical and chemical principles. Bond lengths and angles should match ideal values derived from small molecule crystallography. Atoms should not clash with each other. Residues should adopt backbone conformations that appear frequently in experimentally determined structures. The distribution of hydrophobic and hydrophilic residues should follow patterns observed in folded proteins.

The choice of validation tool matters because different structure types have different error profiles. An X-ray structure refined against electron density maps has been constrained by experimental data. Its validation should focus on whether the final model violates geometric expectations while still fitting the observed density. A predicted structure has no experimental data to constrain it. Its validation must assess whether the model resembles known protein structures in statistically meaningful ways. Applying the wrong validation tool can produce misleading conclusions about structure quality.

## MolProbity: Geometric Validation for Experimental Structures

MolProbity provides all-atom contact analysis and geometric validation for macromolecular structures. The tool evaluates structures against criteria derived from high-resolution experimentally determined proteins. Its core function is identifying problems that would affect the reliability of a deposited or published structure.

### Underlying Principles

MolProbity operates on the principle that protein structures obey geometric rules established by the physics of atomic interactions. Carbon atoms have characteristic van der Waals radii. Peptide bonds are planar. Side chain torsion angles cluster in preferred rotamer conformations. Ramachandran plots show that backbone phi and psi angles occupy restricted regions of conformational space.

The tool performs several distinct analyses. Clash analysis identifies pairs of atoms that are closer than their van der Waals radii permit. The clashscore reports the number of serious steric overlaps per thousand atoms. Rotamer analysis evaluates whether side chain conformations match the most frequently observed conformations in high-resolution structures. Ramachandran analysis classifies backbone angles as favored, allowed, or outlier regions. MolProbity also evaluates bond lengths, bond angles, and the geometry of specific functional groups such as the peptide bond plane.

The MolProbity score combines these individual metrics into a single numerical value. The score is normalized so that it can be compared across structures of different sizes. Lower scores indicate better geometry. The score is calibrated against structures of known resolution, allowing researchers to assess whether their structure's geometry is appropriate for its reported resolution.

### Input Requirements and Usage

MolProbity accepts coordinate files in standard formats including PDB and mmCIF. The tool requires a complete structure with all non-hydrogen atoms assigned. Hydrogen atoms can be added by the tool itself, which is necessary for accurate clash detection. The input structure should have been refined with appropriate geometry restraints, as MolProbity will identify any residual problems.

The tool is available through the MolProbity web server and as part of the Phenix crystallography software suite. The web interface allows users to upload a structure file and receive a comprehensive validation report. The report includes the clashscore, MolProbity score, Ramachandran statistics, rotamer statistics, and detailed lists of problem residues.

### Interpreting MolProbity Output

MolProbity output must be interpreted in the context of the structure's resolution. A structure determined at 2.0 Å resolution should have a MolProbity score below 2.0. A structure at 3.0 Å resolution may have a score between 2.5 and 3.5. The score is meaningful only when compared to structures of similar resolution.

The clashscore is reported as the number of clashes per thousand atoms. A clashscore below 5 is generally considered good. Scores above 20 indicate serious problems that likely require manual inspection and rebuilding. The Ramachandran analysis reports the percentage of residues in favored regions, allowed regions, and outliers. High-quality structures typically have more than 95 percent of residues in favored regions and fewer than 1 percent outliers.

Rotamer analysis identifies side chains that adopt unusual conformations. The percentage of rotamer outliers should be low, typically below 1 percent for well-refined structures. Individual rotamer outliers may represent genuine conformational features or model building errors. Each outlier should be examined in the context of the electron density map.

### Application in Published Research

Recent structural studies demonstrate the standard application of MolProbity validation. A study of lung cancer biomarkers used SWISS-MODEL to generate three-dimensional structures of EGFR, ALK, KRAS, and PD-1, then validated the models with MolProbity scores between 0.67 and 2.09, confirming structural reliability for subsequent docking studies [7]. The score range reflects the varying quality of the templates used for each target.

A study of SLC22A transporter proteins associated with polycystic ovary syndrome and metformin response built structural models in SWISS-MODEL and validated them using MolProbity scores below 1.5, along with ERRAT values above 85 percent and PROSA Z-scores between -6.8 and -8.2 [9]. The combination of multiple validation metrics provides stronger evidence of model quality than any single score.

These examples illustrate the standard practice of reporting MolProbity scores alongside other validation metrics in published structural studies. The scores provide reviewers and readers with quantitative evidence that the models are reliable enough to support the study's conclusions.

## QMEAN: Statistical Potential Validation for Predicted Structures

QMEAN, which stands for Qualitative Model Energy Analysis, provides composite scoring for protein structure models. The tool evaluates predicted structures by comparing them to statistical potentials derived from experimentally determined protein structures. It is widely used in the context of homology modeling and structure prediction.

### Underlying Principles

QMEAN is based on the observation that correctly folded proteins exhibit characteristic patterns of atomic interactions, solvent exposure, and secondary structure. These patterns can be captured as statistical potentials, which describe the likelihood of observing particular geometric arrangements in known protein structures.

The QMEAN score combines four statistical potential terms. The C-beta interaction potential evaluates the pairwise distances between C-beta atoms. The all-atom pairwise energy potential assesses interactions between all atom types. The solvation potential evaluates the burial of hydrophobic residues and the exposure of hydrophilic residues. The torsion angle potential assesses the distribution of backbone and side chain torsion angles.

The composite QMEAN score is normalized to a scale where zero represents the average quality of experimental structures and negative values indicate models that deviate from expected patterns. The score is often reported as a Z-score, which describes how many standard deviations the model's score deviates from the mean of experimental structures of similar size.

### QMEAN in the SWISS-MODEL Workflow

QMEAN is integrated into the SWISS-MODEL homology modeling server. When a user submits a target sequence, SWISS-MODEL generates models and automatically calculates QMEAN scores for each model. The server also provides QMEANDisCo, a variant that incorporates distance constraints from homologous structures to improve discrimination between correct and incorrect models.

The QMEAN score is used to select the best model from multiple alternatives. Models with higher QMEAN scores, closer to zero, are preferred over models with lower scores. The score also provides an absolute measure of model quality that can be compared across different targets.

### Interpreting QMEAN Output

QMEAN output includes the composite score, the individual potential terms, and a Z-score. The composite score is typically negative, with values closer to zero indicating better quality. A model with a QMEAN score of -1.0 is generally considered good, while a score of -4.0 or lower indicates poor quality.

The Z-score provides a statistical context for the composite score. A Z-score of -2.0 means the model's score is two standard deviations below the mean of experimental structures. Models with Z-scores below -4.0 are unlikely to represent correctly folded proteins.

The individual potential terms can identify specific problems. A poor solvation potential suggests incorrect burial of hydrophobic residues. A poor torsion angle potential suggests incorrect backbone or side chain conformations. These diagnostic terms help researchers identify which regions of the model may need attention.

### Application in Published Research

A study of L-asparaginase from Streptomyces koyangensis used homology modeling with SWISS-MODEL and AlphaFold2 to generate structures, then validated the enzyme-oncoprotein interactions through molecular docking and molecular dynamics simulations [8]. The study reported binding energy of -13.8 kcal/mol and RMSD below 2.5 Å over 100 nanoseconds of simulation, with MM-PBSA calculations yielding -68.4 ± 5.2 kcal/mol [8]. These results demonstrate the integration of predicted structures into a broader computational validation workflow.

A study of the ASPM protein in hepatocellular carcinoma used structural modeling, molecular docking, 100-nanosecond molecular dynamics simulations, and MM-GBSA binding energy calculations to assess ligand-target interactions [10]. The study identified a lead compound with stable binding within the ASPM calponin domain and favorable pharmacokinetic properties [10]. The computational workflow required validated structural models as the foundation for subsequent analyses.

A multi-epitope vaccine design study against Acinetobacter baumannii used modeling and molecular dynamics-based refinements to obtain the tertiary structure of the vaccine candidate, then evaluated the construct for immunogenicity, physicochemical properties, structure, binding to toll-like receptors, solubility, stability, toxicity, allergenicity, and cross-reactivity [11]. The study emphasized that in vitro and in vivo experimental tests are needed to validate the efficacy of the vaccine candidate [11]. This example illustrates the role of structural validation in a broader computational immunology workflow.

## At a Glance: Tool Selection Decision Table

| Validation Aspect | MolProbity | QMEAN |
|---|---|---|
| Primary structure type | Experimental X-ray crystallography | Predicted models from homology or deep learning |
| Core analysis method | All-atom contacts and geometric criteria | Statistical potentials from known structures |
| Key output metrics | Clashscore, MolProbity score, Ramachandran statistics, rotamer statistics | Composite QMEAN score, Z-score, individual potential terms |
| Input format | PDB or mmCIF coordinate files | Model coordinate files, typically from SWISS-MODEL |
| Resolution dependence | Score interpretation depends on structure resolution | Score normalized against experimental structures |
| Typical use case | Validating refined crystal structures before deposition | Selecting best model from prediction runs |
| Availability | Web server and Phenix integration | SWISS-MODEL server integration |
| Published example | SLC22A transporter models with scores below 1.5 [9] | L-asparaginase models used for docking and MD [8] |

## Core Principles of Structure Validation

### Geometric Criteria for Experimental Structures

Experimental structures must satisfy geometric criteria that reflect the physical chemistry of proteins. Bond lengths and angles should match ideal values derived from small molecule crystallography. The peptide bond should be planar. Side chains should adopt conformations that are energetically favorable.

MolProbity evaluates these criteria systematically. The clashscore identifies atoms that violate van der Waals separation requirements. The Ramachandran analysis identifies backbone conformations that are sterically disallowed. The rotamer analysis identifies side chain conformations that are rarely observed in high-resolution structures.

These geometric criteria are particularly important for X-ray structures because refinement procedures can produce models that fit the electron density while containing subtle geometric distortions. The validation step ensures that the final model is both experimentally supported and physically reasonable.

### Statistical Potentials for Predicted Structures

Predicted structures cannot be validated against experimental data because no such data exists for the target protein. Instead, validation relies on statistical potentials derived from the database of known protein structures. These potentials capture the likelihood of observing particular geometric arrangements in correctly folded proteins.

QMEAN applies these potentials to evaluate whether a predicted model resembles known protein structures. The composite score reflects the overall similarity, while individual potential terms identify specific deviations. This approach is particularly useful for selecting among multiple models generated by different templates or prediction methods.

### The Complementarity of Validation Approaches

MolProbity and QMEAN are complementary instead of competing tools. A researcher may use both in a single workflow. For example, a homology model built with SWISS-MODEL can be evaluated with QMEAN to assess overall quality, then refined against experimental data if such data becomes available, and finally validated with MolProbity to ensure geometric correctness.

The choice of tool depends on the structure's origin and the research question. An X-ray crystallographer validating a newly refined structure should use MolProbity. A computational biologist evaluating a predicted model should use QMEAN. A researcher working with both types of structures should understand both tools and apply them appropriately.

## Practical Workflow for Structure Validation

### Step 1: Determine Structure Origin

The first decision is whether the structure is experimental or predicted. Experimental structures come from X-ray crystallography, cryo-electron microscopy, or NMR spectroscopy. Predicted structures come from homology modeling, threading, or deep learning methods such as AlphaFold.

This determination guides all subsequent validation decisions. Experimental structures require geometric validation against physical criteria. Predicted structures require statistical validation against known protein structures.

### Step 2: Select the Appropriate Validation Tool

For experimental structures, use MolProbity. The tool provides comprehensive geometric validation including clash detection, Ramachandran analysis, and rotamer analysis. The output can be compared to structures of similar resolution to assess whether the geometry is appropriate.

For predicted structures, use QMEAN. The tool provides composite scoring that reflects overall model quality. The output can be used to select among multiple models and to identify specific problem regions.

### Step 3: Prepare the Input Structure

Both tools require coordinate files in standard formats. MolProbity can add hydrogen atoms to the structure, which is necessary for accurate clash detection. QMEAN requires a complete model with all residues assigned.

The input structure should be checked for completeness before validation. Missing residues or atoms can affect the validation results. Some validation tools can handle incomplete structures, but the results should be interpreted with caution.

### Step 4: Run the Validation and Record Results

Run the selected validation tool and record all output metrics. For MolProbity, record the clashscore, MolProbity score, Ramachandran statistics, and rotamer statistics. For QMEAN, record the composite score, Z-score, and individual potential terms.

The results should be recorded in a laboratory notebook or electronic data management system. This documentation is essential for reproducibility and for supporting publication claims.

### Step 5: Interpret Results in Context

Interpret the validation results in the context of the structure's origin and intended use. A MolProbity score of 2.5 may be acceptable for a 3.0 Å resolution structure but poor for a 1.5 Å resolution structure. A QMEAN Z-score of -3.0 may be acceptable for a difficult modeling target but concerning for a straightforward homology model.

The interpretation should also consider the intended use of the structure. A structure used for molecular docking requires higher quality than a structure used for a coarse-grained analysis. The validation results should be sufficient to support the conclusions drawn from the structure.

### Step 6: Document and Report

Document the validation results in the methods section of any publication or report. Include the specific scores and the tool versions used. This documentation allows readers to assess the reliability of the structure and to reproduce the validation.

The reporting should follow community standards. For experimental structures, report the MolProbity score, clashscore, and Ramachandran statistics. For predicted structures, report the QMEAN score and Z-score. Include the resolution for experimental structures and the template information for predicted structures.

## Records and Measurements for Validation

### Essential Records for Experimental Structures

Researchers should maintain a validation record for each experimental structure. The record should include the structure identifier, resolution, R-factor, and free R-factor. The MolProbity validation report should be archived with the structure.

The validation record should document any manual interventions. If a residue was rebuilt to resolve a clash, the intervention should be noted. If a rotamer outlier was retained because it is supported by electron density, this decision should be documented.

### Essential Records for Predicted Structures

Researchers should maintain a validation record for each predicted structure. The record should include the target sequence identifier, template structures used, sequence identity to templates, and the QMEAN score. The model file should be archived with the validation report.

The validation record should document the model selection process. If multiple models were generated, the scores for each model should be recorded. The rationale for selecting the final model should be documented.

### Comparative Measurements Across Structures

Validation scores become more meaningful when compared across related structures. A researcher studying a protein family can compare MolProbity scores across family members to identify structures that may have quality issues. A researcher evaluating multiple prediction methods can compare QMEAN scores to identify the most reliable approach.

These comparative measurements should be recorded in a structured format that allows statistical analysis. A spreadsheet or database with columns for structure identifier, validation tool, score, and relevant metadata provides a useful framework.

## Common Failure Patterns in Structure Validation

### Overinterpretation of Single Scores

A common failure is relying on a single validation score to judge structure quality. The MolProbity score combines multiple metrics, but it does not capture all aspects of structure quality. A structure with a good MolProbity score may still have errors in loop regions or side chain placements that are not captured by the composite metric.

Similarly, a QMEAN score provides an overall assessment but does not identify all problems. A model with a good QMEAN score may have local errors in specific regions. The individual potential terms should be examined to identify these local problems.

### Applying the Wrong Tool to a Structure Type

Applying MolProbity to a predicted structure or QMEAN to an experimental structure produces misleading results. MolProbity's geometric criteria are designed for structures refined against experimental data. A predicted structure will have geometric deviations that reflect the modeling process, not experimental error. QMEAN's statistical potentials are designed for predicted structures. An experimental structure may have features that are not captured by the statistical potentials.

Researchers should verify the origin of their structure before selecting a validation tool. If the structure is experimental, use MolProbity. If the structure is predicted, use QMEAN.

### Ignoring Resolution or Template Quality

Validation scores must be interpreted in the context of structure resolution or template quality. A low-resolution experimental structure will have higher MolProbity scores than a high-resolution structure of the same protein. A predicted model built from a distant template will have lower QMEAN scores than a model built from a close homolog.

Researchers should compare their validation scores to appropriate reference values. For experimental structures, compare to structures of similar resolution. For predicted structures, compare to models built from templates of similar quality.

### Failing to Document Validation Decisions

Validation involves judgment calls that should be documented. If a residue with poor geometry is retained because it is supported by experimental data, this decision should be recorded. If a model with a moderate QMEAN score is selected because it is the best available, this rationale should be documented.

Undocumented validation decisions create problems for reproducibility. Other researchers cannot assess the reliability of the structure without knowing how validation results were interpreted.

## Limitations of Validation Tools

### MolProbity Limitations

MolProbity cannot detect all errors in experimental structures. The tool evaluates geometry but does not assess whether the structure fits the experimental data. A structure with excellent geometry may still have errors in the placement of residues relative to the electron density. The validation of data fit requires separate tools that compare the model to the diffraction data.

MolProbity also has limited sensitivity to errors in disordered regions. Residues with weak electron density may be placed incorrectly without producing geometric violations. The validation report may not flag these errors because the geometry appears acceptable.

### QMEAN Limitations

QMEAN cannot detect errors that are consistent with statistical potentials. A predicted model may have a good QMEAN score while containing errors that are not captured by the potential terms. The tool evaluates overall similarity to known structures but does not assess the correctness of specific interactions.

QMEAN is also limited by the diversity of the reference structure database. Proteins with unusual folds or compositions may receive lower scores even when the model is correct. The statistical potentials reflect the distribution of known structures, which may not represent all possible protein folds.

### General Limitations

All validation tools are limited by the quality of the reference data. The geometric criteria in MolProbity are derived from high-resolution experimental structures. The statistical potentials in QMEAN are derived from the database of known structures. Errors in the reference data propagate to the validation tools.

Validation tools also cannot assess the biological relevance of a structure. A structure may pass all validation criteria while representing a conformation that is not biologically meaningful. The validation of biological relevance requires experimental data and functional studies.

## Quality Controls for Structural Studies

### Pre-Deposition Validation for Experimental Structures

Before depositing an experimental structure in the Protein Data Bank, researchers should run comprehensive validation. The validation should include MolProbity analysis as well as checks for data completeness, refinement quality, and agreement with the experimental data.

The Protein Data Bank requires validation reports for all deposited structures. These reports include MolProbity scores and other quality metrics. Researchers should review the validation report before submission to identify and address any problems.

### Model Selection Criteria for Predicted Structures

When generating predicted structures, researchers should use validation scores to select among alternative models. The QMEAN score provides a quantitative basis for model selection. Models with higher scores should be preferred over models with lower scores.

The model selection should also consider the template quality and sequence identity. A model built from a close homolog with a high QMEAN score is more reliable than a model built from a distant homolog with a similar score. The template information should be documented with the model.

### Integration with Experimental Validation

Predicted structures should be validated experimentally when possible. The computational predictions can be tested through mutagenesis, binding assays, or structural determination. The experimental results provide the ultimate validation of the predicted structure.

The integration of computational and experimental validation is particularly important for studies that use predicted structures to support biological conclusions. The computational validation provides initial confidence, but experimental confirmation is required for strong claims.

## Safety and Regulatory Context

### Data Management and Reproducibility

Structural validation generates data that should be managed according to reproducible research practices. The validation inputs, tool versions, and outputs should be documented to allow others to reproduce the analysis. The documentation should follow the standards of the research community.

Training resources are available for researchers who need to develop their computational skills. The Galaxy Training Network provides accessible workflow training and analysis tutorials for bioinformatics applications [4]. The Carpentries offers foundational computing, data, shell, Git, and programming training [6]. These resources support the development of reproducible analysis workflows.

### Community Standards for Structure Reporting

The structural biology community has established standards for reporting validation results. These standards ensure that structures can be compared across studies and that quality assessments are consistent. Researchers should follow these standards when reporting validation results in publications.

The standards include reporting the specific validation tool versions, the scores obtained, and the context for interpreting the scores. For experimental structures, the resolution and refinement statistics should be reported alongside the validation scores. For predicted structures, the template information and modeling method should be reported.

### Professional Escalation Criteria

Researchers should escalate validation concerns to appropriate professionals when certain criteria are met. If a structure fails validation criteria and the problems cannot be resolved through rebuilding, the structure should be reviewed by a crystallographer or structural biologist with appropriate expertise.

If a predicted structure has poor validation scores and no better model can be generated, the limitations should be acknowledged in any publication or report. The structure should not be used to support conclusions that require high structural confidence.

## Professional Escalation Criteria for Validation Concerns

### When to Consult a Crystallographer

A crystallographer should be consulted when an experimental structure has persistent validation problems. These problems include high clashscores that cannot be resolved through rebuilding, Ramachandran outliers that are not supported by electron density, and rotamer outliers that cannot be explained by the structural context.

The crystallographer can review the refinement strategy, examine the electron density maps, and recommend appropriate corrections. The consultation should occur before the structure is deposited or published.

### When to Consult a Computational Biologist

A computational biologist should be consulted when a predicted structure has poor validation scores. The consultant can evaluate the modeling strategy, identify alternative templates, and recommend improved modeling approaches.

The consultation should occur when the QMEAN score is significantly worse than expected for the target and template quality. The consultant can help determine whether the poor score reflects a genuine modeling problem or an inherent difficulty of the target.

### When to Acknowledge Limitations

Limitations should be acknowledged when validation scores indicate that a structure is not reliable enough to support the intended conclusions. The acknowledgment should be explicit in any publication or report that uses the structure.

The acknowledgment should describe the validation results and explain why the structure was used despite the limitations. The explanation should be honest about the potential impact of the limitations on the study's conclusions.

## Building a Validation Decision Matrix for Mixed Structure Workflows

Researchers often encounter a practical problem that standard tool descriptions do not address: how to handle projects that contain both experimental and predicted structures, or structures that transition between categories during the research process. A homology model refined against low-resolution electron density, an AlphaFold prediction used as a search model for molecular replacement, or a comparative study that mixes PDB-deposited crystal structures with newly predicted models all require a coherent validation strategy. A structured decision matrix helps researchers apply the correct tool at each stage and interpret scores consistently across heterogeneous structure sets.

### Defining the Structure Origin Categories

The first step in building a decision matrix is classifying each structure into one of four categories. The first category is experimentally determined and refined, which includes X-ray crystal structures that have been through refinement against diffraction data. The second category is experimentally determined but not fully refined, which includes cryo-electron microscopy maps with preliminary models or molecular replacement solutions that have not completed refinement. The third category is predicted from a close template, defined as sequence identity above 30 percent to a template with known structure. The fourth category is predicted from a distant template or ab initio, which includes models built from templates below 30 percent sequence identity or generated without a template.

This classification matters because the validation approach differs for each category. A fully refined experimental structure requires MolProbity as the primary validation tool. A preliminary experimental model may benefit from both MolProbity for geometry and a density fit assessment. A close-template prediction should use QMEAN as the primary tool, with MolProbity as a secondary geometric check. A distant-template prediction requires QMEAN with careful interpretation and additional checks from complementary tools.

### Applying the Decision Matrix

For each structure in a mixed workflow, researchers should apply the following decision logic. If the structure is experimental and refined, run MolProbity and compare scores against resolution-matched reference structures. If the structure is experimental but preliminary, run MolProbity to identify geometric problems that need correction during refinement, and assess density fit separately. If the structure is predicted from a close template, run QMEAN for model selection and overall quality assessment, then run MolProbity to identify local geometric issues that may affect downstream applications such as docking. If the structure is predicted from a distant template, run QMEAN and interpret the Z-score with caution, recognizing that lower scores may reflect template limitations instead of modeling errors.

The decision matrix also guides the interpretation of conflicting results. A predicted structure with a good QMEAN score but poor MolProbity geometry requires investigation. The poor geometry may indicate local errors that QMEAN does not capture, or it may reflect the modeling method's tendency to produce strained conformations. A predicted structure with poor QMEAN but good MolProbity geometry may indicate that the model is geometrically plausible but does not resemble known folds, which is a concern for distant-template models.

### Recording Validation Decisions in a Structured Format

A spreadsheet or database with one row per structure provides a practical record system for mixed workflows. Each row should include the structure identifier, origin category, validation tool applied, primary score, secondary scores, resolution or template information, and the date of validation. The record should also include a column for the validation decision, such as accepted, accepted with caveats, or rejected.

The record system should capture the rationale for each decision. If a structure with a borderline QMEAN score was accepted because it was the best available model for the target, this rationale should be documented. If a structure with a high clashscore was accepted because the clashes were localized to a flexible loop region, this should be noted. This documentation supports reproducibility and provides a basis for revisiting validation decisions if new information becomes available.

### Troubleshooting Common Mixed-Workflow Problems

One common problem is comparing validation scores across structures of different origins. A MolProbity score of 2.5 for an experimental structure at 3.0 Å resolution is not directly comparable to a MolProbity score of 2.5 for a predicted structure. The decision matrix should include a normalization step that converts scores to a common scale or that flags comparisons across categories as invalid.

Another common problem is using a predicted structure as a search model for molecular replacement without revalidating the resulting experimental structure. The molecular replacement solution should be treated as a preliminary experimental structure and validated with MolProbity after refinement. The QMEAN score from the original prediction is no longer the primary validation metric once experimental data has been incorporated.

A third problem is the tendency to report only the most favorable validation score for a structure. If a structure was validated with both QMEAN and MolProbity, both scores should be reported. Selective reporting of favorable scores undermines the credibility of the validation process and prevents other researchers from assessing the structure's reliability.

### Integrating the Decision Matrix with Published Workflows

Published studies demonstrate how the decision matrix operates in practice. The lung cancer biomarker study used SWISS-MODEL to generate structures and validated them with MolProbity scores between 0.67 and 2.09 [7]. These structures fall into the close-template prediction category, and the reported MolProbity scores provide the secondary geometric check recommended by the decision matrix. The SLC22A transporter study similarly used SWISS-MODEL and reported MolProbity scores below 1.5 alongside ERRAT and PROSA Z-scores [9], demonstrating the multi-metric approach that the decision matrix supports.

The L-asparaginase study used both SWISS-MODEL and AlphaFold2 for homology modeling [8], which places the structures in the predicted category. The study's validation focused on downstream analyses such as docking and molecular dynamics instead of reporting QMEAN scores directly [8]. This approach is consistent with the decision matrix, which recommends QMEAN for model selection and then focuses on functional validation for downstream applications.

### When to Escalate Validation Concerns

The decision matrix includes explicit escalation criteria. If a predicted structure has a QMEAN Z-score below -4.0, the model should not be used for downstream applications without consulting a computational biologist. If an experimental structure has a clashscore above 20 or more than 5 percent Ramachandran outliers, a crystallographer should review the structure before deposition or publication.

Escalation is also appropriate when validation results conflict with biological expectations. If a predicted structure passes all validation criteria but places a known functional residue in an implausible position, the structure should be reviewed before use. Validation tools assess geometric and statistical quality, but they do not assess biological plausibility. The decision matrix should include a final check for consistency with known functional and biochemical data.

### Training and Reproducibility Considerations

Researchers implementing a decision matrix should ensure that all team members understand the validation tools and the interpretation of scores. Training resources are available through the Galaxy Training Network, which provides accessible workflow training and analysis tutorials [4], and The Carpentries, which offers foundational computing and data skills [5]. These resources support the development of reproducible validation workflows.

The decision matrix itself should be documented as part of the research protocol. The documentation should include the classification criteria, the validation tools applied to each category, the score thresholds used for decisions, and the escalation criteria. This documentation allows other researchers to understand and reproduce the validation process.

### Limitations of the Decision Matrix Approach

The decision matrix simplifies a complex decision process, and researchers should recognize its limitations. The category boundaries are not always clear. A structure refined at very low resolution may behave more like a predicted structure than an experimental structure. A prediction built from a template at 29 percent sequence identity may be more reliable than one built from a template at 31 percent identity, depending on the structural conservation of the target.

The score thresholds in the decision matrix are starting points, not absolute rules. A structure with a QMEAN Z-score of -3.5 may be acceptable for a difficult target with no close homologs, while a structure with a Z-score of -2.5 may be unacceptable for a target with many available templates. The decision matrix should be applied with judgment and documented rationale.

## Frequently Asked Questions

### What is the main difference between MolProbity and QMEAN?

MolProbity validates experimental structures by evaluating geometric criteria such as atomic clashes, Ramachandran angles, and rotamer conformations against ideal values derived from high-resolution crystal structures. QMEAN validates predicted structures by comparing them to statistical potentials derived from known protein structures. The choice between them depends on whether your structure was determined experimentally or generated by prediction methods.

### Can I use MolProbity to validate a predicted structure?

MolProbity can be run on any coordinate file, but its geometric criteria are designed for structures refined against experimental data. A predicted structure will have geometric deviations that reflect the modeling process instead of experimental error. The MolProbity score may be misleading for predicted structures. QMEAN is the appropriate validation tool for predicted models.

### Can I use QMEAN to validate an experimental structure?

QMEAN can be run on experimental structures, but its statistical potentials are designed to evaluate predicted models. An experimental structure refined against diffraction data may have features that are not captured by the statistical potentials. MolProbity is the appropriate validation tool for experimental structures.

### What is a good MolProbity score?

A good MolProbity score depends on the resolution of the structure. A structure determined at 2.0 Å resolution should have a MolProbity score below 2.0. A structure at 3.0 Å resolution may have a score between 2.5 and 3.5. The score should be compared to structures of similar resolution. Published studies have reported MolProbity scores between 0.67 and 2.09 for models used in downstream analyses [7].

### What is a good QMEAN score?

QMEAN scores are typically negative, with values closer to zero indicating better quality. A score of -1.0 is generally considered good, while a score of -4.0 or lower indicates poor quality. The Z-score provides additional context by describing how many standard deviations the model deviates from the mean of experimental structures.

### Why do published studies report multiple validation metrics?

Published studies often report multiple validation metrics because no single metric captures all aspects of structure quality. A study of SLC22A transporter proteins reported MolProbity scores below 1.5, ERRAT values above 85 percent, and PROSA Z-scores between -6.8 and -8.2 [9]. The combination of metrics provides stronger evidence of model quality than any single score.

### How should I report validation results in a publication?

Report the specific validation tool and version used, the scores obtained, and the context for interpreting the scores. For experimental structures, report the resolution and refinement statistics alongside the MolProbity scores. For predicted structures, report the template information and modeling method alongside the QMEAN score. This documentation allows readers to assess the reliability of the structure.

### What should I do if my structure fails validation?

If an experimental structure fails MolProbity validation, review the specific problems identified and attempt to rebuild the problematic regions. Consult a crystallographer if the problems persist. If a predicted structure fails QMEAN validation, consider alternative templates or modeling methods. Consult a computational biologist if the problems persist. Acknowledge the limitations in any publication or report that uses the structure.

## Related Bioinformatics Guides

- [RNA-Seq Alignment: Choosing the Right Tool and Parameters](/knowledge/bioinformatics/rna-seq-alignment-choosing-the-right-tool-and-parameters)
- [Metagenomics vs Metabarcoding: Choosing the Right Approach for Your Study](/knowledge/bioinformatics/metagenomics-vs-metabarcoding-choosing-the-right-approach-for-your-study)
- [Metagenomics vs Metatranscriptomics: Choosing the Right Approach for Functional Profiling](/knowledge/bioinformatics/metagenomics-vs-metatranscriptomics-choosing-the-right-approach-for-functional-profiling)
- [Spatial Proteomics vs. Single-Cell Proteomics: Choosing the Right Approach](/knowledge/bioinformatics/spatial-proteomics-vs-single-cell-proteomics-choosing-the-right-approach)
- [Gene Set Enrichment Analysis Tools: Choosing the Right One](/knowledge/bioinformatics/gene-set-enrichment-analysis-tools-choosing-the-right-one)

## References and Further Reading

- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Bioconductor](https://bioconductor.org/). Bioconductor Project.
- [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.
- [nf-core Documentation](https://nf-co.re/docs). nf-core.
- [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.
- [Integrative profiling of lung cancer biomarkers EGFR, ALK, KRAS, and PD-1 with emphasis on nanomaterials-assisted immunomodulation and targeted therapy.](https://doi.org/10.3389/fimmu.2025.1649445). 2025.
- [Streptomyces koyangensis L-asparaginase: computational prediction of dual-mechanism BCL-2 interaction in acute lymphoblastic leukemia.](https://doi.org/10.1038/s41598-026-42798-0). 2026.
- [Identification of high-risk SNPs in SLC22A transporter genes: their potential role in PCOS and metformin uptake.](https://doi.org/10.1186/s12863-026-01424-8). 2026.
- [Comparative transcriptomics and computational drug discovery identify ASPM as a key oncogenic driver and therapeutic target in hepatocellular carcinoma.](https://doi.org/10.3389/fbinf.2026.1795889). 2026.
- [Design a multi-epitope vaccine candidate against Acinetobacter baumannii using advanced computational methods.](https://doi.org/10.1186/s13568-025-01913-6). 2025.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.