# What Is a Good Docking Score for Protein-Protein Complexes? Interpreting Scoring Functions and Confidence Metrics

Protein-protein docking generates a numerical score for each predicted binding pose, but researchers frequently misinterpret this value as a binding free energy or as an absolute measure of complex quality. A docking score is a relative ranking metric produced by a scoring function that approximates the physical and chemical complementarity between two protein surfaces. There is no universal threshold that separates good from bad docking scores across all systems, because each scoring function uses different energy terms, units, and calibration data. The practical question is whether the score reliably ranks the correct pose above decoy poses within a given docking run and whether the top-ranked model is consistent with independent biological evidence.

This article provides a framework for interpreting docking scores in protein-protein interaction studies. It covers the meaning of scoring functions, the difference between docking scores and binding affinities, benchmark practices, common failure patterns, and the records and quality controls that make docking results reproducible and defensible in a research report.

## The Role of Scoring Functions in Protein-Protein Docking

Protein-protein docking is a computational procedure that predicts the three-dimensional structure of a complex formed by two or more proteins. The process has two major phases: sampling of possible binding modes between the interacting molecules and scoring for the identification of the correct orientations. Sampling generates many candidate poses, and scoring evaluates each pose to rank them. The scoring phase produces the docking score that researchers must interpret.

Scoring functions are mathematical approximations of the favorability of a protein-protein interaction. They typically combine terms that account for shape complementarity, electrostatic interactions, desolvation effects, and van der Waals forces. For example, the pyDock approach uses van der Waals, electrostatics, and desolvation energy to score docking poses generated by sampling methods such as FTDock or ZDOCK. The choice of energy terms and their weighting determines what the score means and how it should be interpreted.

A docking score is not a measured physical quantity. It is a computed value that reflects how well a particular pose satisfies the assumptions built into the scoring function. Different scoring functions will assign different scores to the same pose, and the same scoring function may rank poses differently for different protein pairs. This means that a score of negative 50 from one program is not comparable to a score of negative 50 from another program, and neither value can be converted to a binding free energy without additional calibration.

The primary use of docking scores is ranking. Within a single docking run, the score orders the sampled poses from most to least favorable according to the scoring function. The researcher then selects the top-ranked pose or a small set of top-ranked poses for further analysis. The score itself does not tell the researcher whether the predicted complex is biologically correct. It only tells the researcher which pose the scoring function considers most favorable among the poses that were sampled.

## Why Docking Scores Are Not Binding Affinities

A common error in docking interpretation is treating the docking score as a prediction of binding affinity. Binding affinity is a thermodynamic quantity that describes the strength of an interaction, typically expressed as a dissociation constant or a free energy of binding. Docking scores are approximations that may correlate with affinity in some cases, but they are not calibrated to predict affinity values.

The scoring functions used in docking are designed to discriminate near-native poses from incorrect poses, not to reproduce absolute binding energies. The energy terms are simplified and often ignore entropic contributions, explicit water molecules, conformational changes upon binding, and other factors that contribute to real binding free energies. As a result, a docking score that looks favorable does not guarantee that the complex forms in solution, and a docking score that looks unfavorable does not prove that the complex cannot form.

Some docking protocols include a separate affinity estimation step. For example, pyDock offers modules for binding affinity estimation in addition to binding mode modeling and interface prediction. These affinity estimates use different methods and should be evaluated separately from the docking score used for pose ranking. Researchers should report both the ranking score and any affinity estimate, but they should not conflate the two.

The practical implication is that docking scores should be used to compare poses within a docking run, not to compare different protein pairs or to make claims about binding strength. If the research question requires binding affinity values, the researcher should use experimental methods or a validated affinity prediction tool, and the docking score should be treated as a preliminary filter instead of a final answer.

## At a Glance: Interpreting Docking Scores by Context

The table below summarizes how docking scores should be interpreted depending on the research context. The same numerical score can have different meaning depending on the scoring function, the sampling method, and the question being asked.

| Research Context | What the Score Indicates | Practical Interpretation | Common Pitfall |
| --- | --- | --- | --- |
| Pose ranking within one docking run | Relative favorability of sampled poses | Select top-ranked pose or top cluster for analysis | Treating the score as an absolute energy threshold |
| Comparison across different protein pairs | Little to no meaning | Do not compare scores between different systems | Claiming one complex is stronger because its score is lower |
| Benchmark validation of a docking protocol | Performance of the scoring function on known complexes | Compare rank of near-native poses against decoys | Assuming good benchmark performance guarantees success on new targets |
| Affinity prediction | Approximate correlation only | Use dedicated affinity estimation tools or experiments | Reporting docking score as a binding free energy |
| Virtual screening of interaction partners | Relative ranking of candidate partners | Use score to prioritize candidates for experimental testing | Selecting candidates based on score alone without biological filters |

The central rule is that docking scores are ranking tools within a defined search space. Their reliability depends on the quality of sampling, the appropriateness of the scoring function for the system, and the availability of independent evidence to support the prediction.

## Core Principles for Interpreting Docking Scores

### Scoring Functions Are System Dependent

The performance of a scoring function varies by protein type, interaction mode, and structural flexibility. A scoring function that works well for rigid globular proteins may perform poorly for complexes involving disordered regions or large conformational changes. Researchers should know which scoring function was used and what assumptions it makes before interpreting its output.

The pyDock scoring scheme, which combines van der Waals, electrostatics, and desolvation energy, has shown consistently good prediction performance in community-wide assessment experiments such as CAPRI and CASP. This track record means the method is useful for many protein-protein systems, but it does not mean every prediction is correct. The scoring function is one component of a pipeline that also includes sampling, filtering, and validation.

### Sampling Quality Limits Score Interpretation

A docking score can only rank the poses that were generated during sampling. If the sampling phase misses the correct binding mode, then no score can identify it. Poor sampling can result from insufficient rotational search, incorrect treatment of flexibility, or a starting structure that differs from the biologically relevant conformation.

Researchers should examine the distribution of scores across all sampled poses, beyond the top hit. A large gap between the top score and the next best score may indicate a clear prediction, while a dense cluster of similar scores suggests ambiguity. The number of poses sampled and the diversity of the top-ranked poses are important context for interpreting any single score.

### Interface Prediction Provides Supporting Evidence

Docking scores are more convincing when supported by interface analysis. Prediction of interface and hot-spot residues is useful for guiding and interpreting mutagenesis experiments and for understanding functional and mechanistic aspects of the interaction. If the top-ranked docking pose places known functional residues at the interface, the score gains credibility. If the predicted interface contradicts experimental data, the score should be treated with caution regardless of its numerical value.

For example, in a study of the SARS-CoV-2 membrane protein, researchers identified common interacting residues including Phe103, Arg107, Met109, Trp110, Arg131, and Glu135 in the C-terminal domain of the M protein in membrane-spike and membrane-nucleocapsid protein complexes. These residue-level observations provided biological context for the docking predictions and helped interpret the modeled interactions.

### Experimental Restraints Improve Score Reliability

Docking accuracy improves when experimental information is used to guide or filter the scoring. Restraints based on mutagenesis data, cross-linking experiments, or low-resolution structural data can eliminate poses that are inconsistent with known biology. pyDock supports the use of restraints based on experimental data and the inclusion of low-resolution structural data, which narrows the search space and makes the scoring step more meaningful.

Researchers who have experimental information about the interaction should incorporate it into the docking protocol instead of relying on the raw score alone. The score then ranks poses within the set that satisfies the experimental constraints, which is a more interpretable result.

## Practical Workflow for Docking Score Interpretation

### Step 1: Define the Biological Question

Before running docking, specify what the prediction is meant to answer. Possible questions include identifying the likely binding mode of two known partners, screening a set of candidate partners, interpreting the effect of a mutation on binding, or generating a structural model for drug discovery. The question determines which scoring function, sampling strategy, and validation approach are appropriate.

For biomedical applications, docking can help interpret pathological mutations involved in protein-protein interactions or provide modeled structural data for drug discovery targeting protein-protein interactions. The intended use of the model should be documented before the docking run begins.

### Step 2: Prepare Input Structures

The quality of the input structures directly affects the reliability of the docking score. Structures from experimental determination are preferred, but when experimental structures are unavailable, modeled structures can be used with appropriate caution. Template-free modeling approaches may be necessary when suitable homologous templates are lacking.

In the SARS-CoV-2 membrane protein study, the lack of suitable homologous templates with more than 30 percent sequence identity led the researchers to construct models using template-free modeling with Robetta and trRosetta servers. The choice of modeling method affected the quality of the starting structure, which in turn affected the docking results. Researchers should validate modeled structures before using them in docking and should report the modeling method and validation metrics.

### Step 3: Select Sampling and Scoring Methods

Choose a docking protocol that matches the system. Rigid-body docking is appropriate for complexes with limited conformational change, while flexible docking or molecular dynamics may be needed for systems with significant rearrangement. The scoring function should be selected based on its performance in benchmarks and its suitability for the interaction type.

Community-wide assessment experiments provide information about which methods perform well on which types of problems. Researchers should consult the literature for benchmark results and should document the version of the software and the parameters used.

### Step 4: Run Docking and Examine Score Distributions

After the docking run, examine the distribution of scores across all generated poses. Look for the following features:

- The range of scores and the gap between the top score and subsequent scores
- The number of poses with scores close to the top score
- The structural diversity of the top-ranked poses
- Whether the top-ranked poses cluster in a similar binding mode

A clear prediction is indicated by a top-ranked pose that is structurally distinct from the rest of the population and separated by a meaningful score gap. An ambiguous prediction is indicated by many poses with similar scores but different binding modes.

### Step 5: Apply Biological Filters

Filter the top-ranked poses using independent biological information. Known interface residues, mutation data, conservation patterns, and functional annotations can all be used to select poses that are biologically plausible. The docking score should be used to rank poses within the biologically filtered set, not to override biological evidence.

### Step 6: Validate with Independent Methods

The most reliable docking predictions are those confirmed by independent computational or experimental methods. Molecular dynamics simulations can assess the stability of the predicted complex. Interface analysis can identify hot-spot residues that are consistent with the predicted binding mode. Experimental validation, such as mutagenesis or binding assays, provides the strongest confirmation.

In the study of amiodarone-induced pulmonary fibrosis, network pharmacology was combined with molecular docking to identify key targets and pathways. The docking results were interpreted within a broader framework of gene ontology and pathway enrichment analysis, providing multiple lines of evidence for the predicted interactions.

### Step 7: Document the Interpretation

Record the docking score, the scoring function, the sampling method, the input structures, and the validation steps. This documentation is essential for reproducibility and for defending the interpretation in a research report. The score should be reported with its context, not as an isolated number.

## Options and Tradeoffs in Scoring Function Selection

### Energy-Based Scoring Functions

Energy-based scoring functions combine physical energy terms such as van der Waals, electrostatics, and desolvation. The pyDock approach is an example of this category. These functions are interpretable because the individual energy terms can be examined to understand why a pose scored well or poorly. The tradeoff is that they may not capture all factors that contribute to binding, and they may require more computational time than simpler functions.

### Knowledge-Based Scoring Functions

Knowledge-based scoring functions derive statistical potentials from known protein structures and interactions. They are fast and can capture patterns that are difficult to model physically. The tradeoff is that they are dependent on the training set and may not generalize to novel interaction types.

### Hybrid and Machine Learning Scoring Functions

Hybrid scoring functions combine multiple energy terms, and machine learning approaches train scoring models on large datasets of known complexes. These methods can achieve good benchmark performance but may be less interpretable than energy-based functions. The researcher may not be able to explain why a particular pose received a particular score.

### Tradeoffs in Practice

The choice of scoring function involves tradeoffs between speed, interpretability, and accuracy. For large-scale screening, faster functions may be necessary even if they are less accurate. For detailed analysis of a single complex, a more rigorous scoring function with interpretable energy terms may be preferable. The researcher should document the rationale for the choice and should acknowledge the limitations of the selected approach.

## Observations and Measurements for Score Interpretation

### Score Distributions

The distribution of scores across all sampled poses is a key observation for interpreting a docking result. A histogram or scatter plot of scores can reveal whether the top score is an outlier or part of a continuous distribution. The shape of the distribution provides information about the difficulty of the docking problem and the confidence that can be placed in the top-ranked pose.

### Rank of Near-Native Poses

In benchmark studies where the correct complex structure is known, the rank of the near-native pose among all scored poses is a useful metric. A near-native pose ranked first indicates that the scoring function successfully identified the correct binding mode. A near-native pose ranked outside the top 10 suggests that the scoring function has difficulty with that type of complex.

### Score Gaps

The difference between the top score and the next best score can indicate the clarity of the prediction. A large gap suggests that the scoring function strongly favors one pose over all others. A small gap suggests that multiple poses are nearly equivalent according to the scoring function, and the prediction should be treated as uncertain.

### Interface Residue Conservation

Conservation of interface residues across related proteins provides supporting evidence for a docking prediction. If the predicted interface includes highly conserved residues that are known to be important for function, the prediction is more credible. If the predicted interface includes variable residues with no known functional role, the prediction should be viewed with caution.

### Comparison with Experimental Data

Any available experimental data should be compared with the docking prediction. Mutagenesis data that identify critical residues, cross-linking data that constrain the distance between residues, and low-resolution structural data can all be used to evaluate the docking score. A docking score that ranks a pose consistent with experimental data is more reliable than a score that ranks a pose contradicting the data.

## Records and Documentation for Docking Studies

### Required Records

A docking study should be documented with sufficient detail that another researcher can reproduce the results. The following records should be maintained:

- Input structure files and their sources, including experimental structures or modeling methods
- Software versions and parameter settings for sampling and scoring
- The number of poses generated and the number retained after filtering
- The scoring function and the individual energy terms for the top-ranked poses
- The distribution of scores across all poses
- Any experimental restraints or filters applied
- The criteria used to select the final model

### Reproducibility Standards

Reproducibility is a core requirement for computational research. Training resources from organizations such as the Galaxy Training Network and nf-core emphasize the importance of documented workflows and version control. Researchers should follow similar standards for docking studies, recording the exact commands and parameters used to produce the results.

### Reporting in Publications

When reporting docking results in a publication, include the following information:

- The docking software and version
- The scoring function and its energy terms
- The input structures and their validation
- The number of poses sampled and scored
- The score of the top-ranked pose and the score distribution
- The criteria for selecting the final model
- The limitations of the scoring approach

This level of detail allows readers to assess the reliability of the prediction and to compare it with other studies.

## Common Failure Patterns in Docking Score Interpretation

### Treating Scores as Absolute Energies

The most common failure is interpreting a docking score as a binding free energy. A score of negative 80 does not mean the binding energy is negative 80 kilocalories per mole. The score is a relative ranking metric that depends on the scoring function and the system. Researchers who report docking scores as energies mislead their readers and may draw incorrect conclusions about binding strength.

### Comparing Scores Across Different Systems

Docking scores from different protein pairs cannot be compared directly. A score that indicates a good prediction for one complex may indicate a poor prediction for another complex. Researchers should never claim that one interaction is stronger than another based solely on docking scores.

### Ignoring Sampling Limitations

A docking score can only rank the poses that were generated. If the sampling missed the correct binding mode, the top score is meaningless. Researchers should examine the diversity of the sampled poses and should consider whether the sampling protocol was adequate for the system.

### Overinterpreting the Top-Ranked Pose

The top-ranked pose is not necessarily the correct structure. Scoring functions are approximations, and the correct pose may be ranked second, fifth, or not sampled at all. Researchers should examine multiple top-ranked poses and should use independent evidence to select the final model.

### Neglecting Biological Context

Docking scores are most useful when interpreted within a biological context. A score that ranks a pose with known functional residues at the interface is more meaningful than a score that ranks a pose with no biological support. Researchers who ignore biological context risk selecting a structurally plausible but biologically irrelevant model.

### Failing to Validate

Docking predictions should be validated with independent methods. Molecular dynamics simulations, interface analysis, and experimental data can all provide validation. Researchers who report docking scores without validation leave their predictions open to question.

## Limitations of Docking Scores

### Conformational Flexibility

Most docking methods treat proteins as rigid or partially flexible. Real protein-protein interactions often involve conformational changes upon binding, including side-chain rearrangements and backbone movements. Scoring functions that do not account for these changes may misrank poses that require conformational adjustment.

### Water Molecules and Ions

Water molecules and ions at the protein-protein interface can play critical roles in binding. Most scoring functions do not explicitly include water molecules, and this omission can affect the accuracy of the score. The desolvation energy term in pyDock partially accounts for the effect of water removal, but it is an approximation.

### Entropic Effects

Binding free energy includes entropic contributions from the loss of translational and rotational freedom and from changes in protein dynamics. Most docking scoring functions do not include these terms. As a result, the score reflects enthalpy-like contributions but not the full thermodynamics of binding.

### Scoring Function Calibration

Scoring functions are calibrated on known complexes, and their performance depends on the similarity between the test system and the calibration set. A scoring function that performs well on enzyme-inhibitor complexes may perform poorly on antibody-antigen complexes or on complexes involving disordered proteins.

### Model Quality Dependence

The quality of the input structures limits the quality of the docking prediction. If the input structures are inaccurate, the docking score cannot compensate. Modeled structures should be validated before use, and the uncertainty in the input structures should be acknowledged in the interpretation.

## Quality Controls for Docking Studies

### Input Structure Validation

Validate all input structures before docking. For experimental structures, check the resolution and the completeness of the structure. For modeled structures, use validation tools to assess the quality of the model. The SARS-CoV-2 membrane protein study used intensive validation of the modeled structures before proceeding to docking and molecular dynamics.

### Benchmark Testing

Test the docking protocol on a set of known complexes before applying it to the target system. Benchmark testing reveals whether the scoring function and sampling method are appropriate for the type of interaction being studied. The results of benchmark testing should be reported alongside the target prediction.

### Multiple Scoring Functions

Using multiple scoring functions provides a cross-check on the prediction. If different scoring functions rank the same pose at the top, the prediction is more reliable. If different scoring functions rank different poses at the top, the prediction is uncertain and should be treated with caution.

### Molecular Dynamics Validation

Molecular dynamics simulations can assess the stability of the predicted complex. A complex that remains stable during simulation is more credible than a complex that dissociates or undergoes large conformational changes. The SARS-CoV-2 membrane protein study performed 100 nanosecond molecular dynamics simulations on the model structures to evaluate their stability.

### Independent Biological Evidence

The strongest validation comes from independent biological evidence. Known interface residues, mutation data, and functional annotations can all support or contradict a docking prediction. Researchers should actively seek such evidence instead of relying on the docking score alone.

## Safety and Regulatory Context for Docking Applications

### Biomedical Applications

Docking predictions used in biomedical research have implications for diagnosis, vaccine design, and drug discovery. The structural characterization of protein-protein interactions at atomic resolution has many applications in biomedicine, from diagnosis and vaccine design to drug discovery. Researchers who use docking for these purposes must recognize that computational predictions are hypotheses that require experimental confirmation.

In the study of thyroid-stimulating hormone receptor in autoimmune thyroid diseases, structural biology analyses including protein-protein docking, small-molecule docking, and normal mode dynamics were used to identify prospective modulators of TSHR. These computational predictions were integrated with transcriptomics and imaging data and validated with immunofluorescence staining in a mouse model. The docking results were one component of a multi-layered investigation, not the sole basis for conclusions.

### Drug Discovery Context

Docking is used in drug discovery to identify potential binding partners and to prioritize candidates for experimental testing. The network pharmacology and molecular docking study of amiodarone-induced pulmonary fibrosis used docking to assess the binding affinity of amiodarone to key hub proteins. The docking results were interpreted within a broader framework of target prediction and pathway analysis.

Researchers who use docking for drug discovery should recognize that docking scores are preliminary filters. A favorable docking score does not guarantee that a compound will bind in a cellular context, and an unfavorable score does not rule out binding. Experimental validation is always required.

### Professional Escalation Criteria

Researchers should escalate docking predictions for experimental validation when the prediction will be used for consequential decisions. The following situations warrant escalation:

- The prediction will guide mutagenesis experiments
- The prediction will be used to prioritize drug candidates
- The prediction will be used to interpret pathological mutations
- The prediction will be included in a regulatory submission
- The prediction contradicts established experimental data

In these situations, the docking score should be presented as a hypothesis-generating result, and experimental confirmation should be obtained before the prediction is used for decision-making.

## Practical Assessment Steps for Docking Score Interpretation

### Step 1: Identify the Scoring Function

Determine which scoring function produced the score. Read the software documentation to understand the energy terms and the units. Record the scoring function name and version.

### Step 2: Examine the Score Distribution

Plot the distribution of scores across all sampled poses. Identify the top score, the median score, and the range. Note whether the top score is an outlier or part of a continuous distribution.

### Step 3: Assess the Score Gap

Calculate the difference between the top score and the second-best score. A large gap suggests a clear prediction. A small gap suggests ambiguity.

### Step 4: Evaluate the Top-Ranked Poses

Examine the top 10 or top 20 poses. Determine whether they cluster in a similar binding mode or represent diverse binding modes. If the top poses are diverse, the prediction is uncertain.

### Step 5: Compare with Biological Data

Compare the predicted interface with known functional residues, mutation data, and conservation patterns. Determine whether the prediction is consistent with biological evidence.

### Step 6: Validate with Independent Methods

Run molecular dynamics simulations on the top-ranked pose to assess stability. Use interface analysis tools to identify hot-spot residues. Compare the prediction with any available experimental data.

### Step 7: Document the Interpretation

Record the score, the context, the validation results, and the limitations. Report the docking score with its interpretation, not as an isolated number.

## A Decision Framework for Docking Score Confidence Levels

Interpreting a docking score requires more than knowing the numerical value. Researchers need a structured way to decide whether a score warrants further investigation, experimental validation, or rejection. A confidence-level framework translates raw scores and their context into actionable decisions. This section provides a practical system for assigning confidence levels to docking predictions based on observable criteria that can be recorded and defended in a research report.

### Defining Confidence Levels for Docking Predictions

Confidence levels provide a common language for describing how much weight a docking prediction should carry. The framework below uses four levels that correspond to increasing amounts of supporting evidence. Each level has specific criteria that must be met before a prediction can be assigned to it.

**Level 1: Exploratory.** The docking score ranks a pose at the top of the list, but there is no independent evidence to support the predicted interface. The score distribution shows many poses with similar values, and the top-ranked poses are structurally diverse. This level is appropriate for hypothesis generation only. The prediction should not be used to guide experiments or to draw conclusions about the interaction.

**Level 2: Suggestive.** The top-ranked pose is clearly separated from the next best pose by a meaningful score gap. The top-ranked poses cluster in a similar binding mode instead of representing diverse orientations. The predicted interface includes residues that are known to be functionally important based on prior experimental data. This level supports limited conclusions and can guide targeted validation experiments.

**Level 3: Substantiated.** The prediction meets all Level 2 criteria and is additionally supported by independent computational validation. Molecular dynamics simulations show that the predicted complex remains stable over the simulation time. Interface analysis identifies hot-spot residues that are consistent with the predicted binding mode. The prediction is consistent with experimental restraints such as mutagenesis data or cross-linking constraints. This level supports use of the model for interpreting functional and mechanistic aspects of the interaction.

**Level 4: Confirmed.** The prediction meets all Level 3 criteria and is supported by direct experimental confirmation. Mutagenesis experiments show that mutations at the predicted interface disrupt the interaction. Binding assays confirm that the two proteins interact with measurable affinity. This level is the standard for predictions that will be used for consequential decisions such as drug discovery targeting or interpretation of pathological mutations.

### Assigning Confidence Levels in Practice

The assignment of a confidence level requires a systematic review of the docking output and the available supporting evidence. The following steps provide a reproducible procedure for assigning confidence levels.

**Step 1: Record the raw score and score distribution.** Document the top score, the second-best score, and the range of scores across all sampled poses. Calculate the score gap as the difference between the top score and the second-best score. Record the number of poses with scores within 10 percent of the top score.

**Step 2: Assess structural clustering.** Superimpose the top 10 ranked poses and measure the root mean square deviation between them. If the top poses cluster with low pairwise deviation, the prediction is structurally consistent. If the top poses are diverse, the prediction is ambiguous regardless of the score values.

**Step 3: Check interface residue consistency.** Compare the predicted interface residues with known functional residues from the literature. Look for overlap with residues identified in mutagenesis studies, conservation analyses, or prior structural studies. Record the number of predicted interface residues that match known functional residues.

**Step 4: Review independent validation results.** Determine whether molecular dynamics simulations, interface prediction tools, or experimental restraints have been applied to the top-ranked pose. Record the results of each validation method.

**Step 5: Assign the confidence level.** Apply the criteria from the confidence level definitions. If the prediction meets all criteria for a given level, assign that level. If the prediction meets some but not all criteria, assign the lower level.

**Step 6: Document the assignment.** Record the evidence that supports the confidence level and the evidence that is missing. This documentation allows other researchers to understand why a particular confidence level was assigned.

### Records and Measurements for Confidence Assignment

The confidence level assignment should be supported by specific records that another researcher can inspect. The following measurements should be recorded for each docking prediction.

**Score distribution statistics.** Record the top score, the median score, the mean score, and the standard deviation of scores across all sampled poses. These statistics provide context for evaluating whether the top score is an outlier or part of a continuous distribution.

**Score gap ratio.** Calculate the ratio of the score gap to the standard deviation of the score distribution. A score gap that is large relative to the distribution spread indicates a clearer prediction than a score gap that is small relative to the spread.

**Cluster analysis results.** Record the number of distinct structural clusters among the top 10 poses and the population of each cluster. A prediction with one dominant cluster is more reliable than a prediction with several equally populated clusters.

**Interface residue overlap.** Record the number of predicted interface residues that match known functional residues and the total number of predicted interface residues. The fraction of matching residues provides a quantitative measure of biological consistency.

**Validation method results.** Record the results of each validation method applied to the top-ranked pose. For molecular dynamics simulations, record the simulation length, the root mean square deviation of the complex over the simulation, and whether the complex remained associated. For interface prediction, record the predicted hot-spot residues and whether they match the docking interface.

### Common Failure Patterns in Confidence Assignment

Several recurring errors undermine the reliability of confidence level assignments. Recognizing these patterns helps researchers avoid them.

**Assigning high confidence based on score alone.** A favorable numerical score does not justify a high confidence level. The score is only one component of the evidence. Researchers who assign confidence based on the score without examining the score distribution, structural clustering, or biological consistency produce unreliable assessments.

**Ignoring the score distribution.** The top score is meaningless without knowledge of the distribution of all scores. A top score that is part of a continuous distribution of similar scores indicates ambiguity. A top score that is clearly separated from the rest indicates a more definitive prediction. Researchers who report only the top score omit essential context.

**Overweighting a single validation method.** One validation method is rarely sufficient to substantiate a prediction. A complex that remains stable in a molecular dynamics simulation may still be biologically irrelevant if the predicted interface contradicts known functional data. Confidence should be based on convergence of multiple independent lines of evidence.

**Confusing structural plausibility with biological relevance.** A docking pose can be structurally plausible, with good shape complementarity and favorable energy terms, while being biologically irrelevant because it places known functional residues away from the interface. The confidence level should reflect biological consistency, beyond structural quality.

**Failing to update confidence when new evidence appears.** Confidence levels are not permanent assignments. When new experimental data become available, the confidence level should be reassessed. A prediction that was substantiated by computational validation may be downgraded if new mutagenesis data contradict the predicted interface.

### Troubleshooting Low Confidence Predictions

When a docking prediction receives a low confidence level, the researcher should diagnose the cause before deciding whether to improve the prediction or abandon it. The following troubleshooting steps address common causes of low confidence.

**Check sampling completeness.** If the score distribution shows a continuous range of similar scores with no clear gap, the sampling may have missed the correct binding mode. Increase the number of poses sampled, use a different sampling method, or incorporate experimental restraints to narrow the search space.

**Check input structure quality.** If the input structures are modeled instead of experimental, the model quality may limit the docking accuracy. Validate the input structures and consider whether better templates or alternative modeling methods would improve the starting structures. The SARS-CoV-2 membrane protein study demonstrated that the choice of modeling method affected model quality, with trRosetta providing a better model than Robetta or I-TASSER based on TM-score and RMSD comparisons.

**Check scoring function suitability.** If the scoring function performs poorly on the type of interaction being studied, the scores may not rank the correct pose at the top. Consult benchmark results from community-wide assessments such as CAPRI and CASP to determine whether the scoring function is appropriate for the system.

**Check biological consistency.** If the predicted interface does not overlap with known functional residues, the prediction may be structurally plausible but biologically irrelevant. Re-examine the biological evidence and consider whether the docking protocol should be modified to incorporate restraints based on known functional data.

**Consider alternative binding modes.** If the top-ranked pose has low confidence, examine lower-ranked poses that may be more biologically consistent. The correct binding mode may be ranked second or third by the scoring function. The pyDock approach has shown consistently good prediction performance in community-wide assessments, but no scoring function is perfect for every system.

### Escalation Criteria for Experimental Validation

The confidence level framework provides clear criteria for when a docking prediction should be escalated for experimental validation. The following situations warrant escalation.

**The prediction will guide mutagenesis experiments.** If the predicted interface will be used to select residues for mutation, the prediction should be substantiated or confirmed. Mutagenesis experiments are expensive and time-consuming, and they should not be guided by exploratory predictions.

**The prediction will be used to prioritize drug candidates.** Docking predictions used for drug discovery targeting protein-protein interactions should meet the substantiated level at minimum. The network pharmacology and molecular docking study of amiodarone-induced pulmonary fibrosis used docking to assess binding affinity of amiodarone to key hub proteins, and the results were interpreted within a broader framework of target prediction and pathway analysis.

**The prediction will be used to interpret pathological mutations.** Docking can help interpret pathological mutations involved in protein-protein interactions. Predictions used for this purpose should be substantiated or confirmed because they may inform clinical decisions.

**The prediction contradicts established experimental data.** If the docking prediction conflicts with known experimental results, the discrepancy should be resolved before the prediction is used. The prediction may be incorrect, or the experimental data may need reinterpretation.

**The prediction will be included in a regulatory submission.** Any prediction used in a regulatory context should meet the confirmed level with direct experimental validation.

### Integration with Reproducible Workflow Standards

The confidence level framework should be integrated into a documented workflow that follows reproducibility standards. Training resources from the Galaxy Training Network and nf-core emphasize the importance of documented workflows and version control. The confidence assignment should be recorded as part of the workflow documentation, including the criteria applied and the evidence reviewed.

The Carpentries lessons provide foundational training in computing and data skills that support reproducible research practices. Researchers who apply these practices to docking studies can produce results that are transparent and defensible. The confidence level assignment is a key component of this transparency because it communicates the strength of the evidence supporting each prediction.

The confidence level framework also supports comparison across studies. When researchers report confidence levels alongside docking scores, readers can assess the reliability of predictions from different studies. This comparability is important for building a body of evidence about protein-protein interactions and for identifying predictions that warrant experimental follow-up.

## Frequently Asked Questions

### What Is a Good Docking Score?

There is no universal good docking score. A good score is one that ranks the correct binding mode at the top of the pose list within a given docking run. The numerical value depends on the scoring function, the system, and the sampling protocol. Researchers should evaluate the score in the context of the score distribution and the biological evidence.

### Can Docking Scores Be Compared Between Different Programs?

Docking scores from different programs cannot be compared directly. Each program uses a different scoring function with different energy terms and units. A score from one program has no numerical relationship to a score from another program.

### How Many Top Poses Should Be Examined?

The number of top poses to examine depends on the score distribution and the diversity of the poses. If the top score is clearly separated from the rest, examining the top pose may be sufficient. If multiple poses have similar scores, examining the top 10 or top 20 poses is advisable.

### What Is the Difference Between a Docking Score and a Binding Affinity?

A docking score is a relative ranking metric produced by a scoring function. A binding affinity is a thermodynamic quantity that describes the strength of an interaction. Docking scores are not calibrated to predict binding affinities, and they should not be reported as energies.

### How Can Docking Predictions Be Validated?

Docking predictions can be validated with molecular dynamics simulations, interface analysis, comparison with experimental data, and benchmark testing. The strongest validation comes from experimental confirmation of the predicted interaction.

### What Should Be Reported in a Publication?

A publication should report the docking software and version, the scoring function, the input structures, the number of poses sampled, the score distribution, the criteria for selecting the final model, and the validation results. This information allows readers to assess the reliability of the prediction.

### When Should Docking Scores Not Be Used?

Docking scores should not be used to compare binding strengths between different protein pairs, to predict binding affinities, or to make consequential decisions without experimental validation. They should not be used when the sampling protocol is inadequate for the system or when the input structures are of poor quality.

### How Do Community-Wide Assessments Help With Score Interpretation?

Community-wide assessment experiments such as CAPRI and CASP evaluate the performance of docking methods on blinded test cases. The results provide information about which methods perform well on which types of problems. Researchers can use this information to select appropriate methods and to calibrate their expectations for prediction accuracy.

## Related Bioinformatics Guides

- [Lipidomic Analysis: A Beginner's Guide to Workflows and Data Interpretation](/knowledge/bioinformatics/lipidomic-analysis-a-beginner-s-guide-to-workflows-and-data-interpretation)
- [Pathway Enrichment Analysis for Proteomics: Tools and Interpretation](/knowledge/bioinformatics/pathway-enrichment-analysis-for-proteomics-tools-and-interpretation)
- [Proteomics Analysis Tools: A Comparative Guide for Functional Interpretation](/knowledge/bioinformatics/proteomics-analysis-tools-a-comparative-guide-for-functional-interpretation)
- [Genomic Data Analysis Tools: A Comparative Guide for Researchers](/knowledge/bioinformatics/genomic-data-analysis-tools-a-comparative-guide-for-researchers)
- [How to Interpret Gene Set Enrichment Analysis Results](/knowledge/bioinformatics/how-to-interpret-gene-set-enrichment-analysis-results)

## References and Further Reading

- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Bioconductor](https://bioconductor.org/). Bioconductor Project.
- [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.
- [nf-core Documentation](https://nf-co.re/docs). nf-core.
- [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.
- [Structural Prediction of Protein-Protein Interactions by Docking: Application to Biomedical Problems.](https://pubmed.ncbi.nlm.nih.gov/29412997). Advances in protein chemistry and structural biology, 2018.
- [Thyroid-stimulating hormone receptor mediates peripheral-central neuroimmune crosstalk in autoimmune thyroid diseases.](https://pubmed.ncbi.nlm.nih.gov/42204533). BMC medicine, 2026.
- [Network pharmacology and molecular docking reveal mechanisms of amiodarone-induced pulmonary fibrosis.](https://pubmed.ncbi.nlm.nih.gov/41492010). Scientific reports, 2026.
- [Modeling of Protein Complexes and Molecular Assemblies with pyDock.](https://pubmed.ncbi.nlm.nih.gov/32621225). Methods in molecular biology (Clifton, N.J.), 2020.
- [Structure and dynamics of membrane protein in SARS-CoV-2.](https://pubmed.ncbi.nlm.nih.gov/33353499). Journal of biomolecular structure & dynamics, 2022.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.