# How to Interpret Docking Scores: A Guide to Understanding Binding Affinity Predictions

Molecular docking produces numerical scores that estimate how favorably a small molecule may bind to a protein target. These scores are widely used in early-stage drug discovery, virtual screening, and hit-to-lead optimization, yet many researchers receive docking outputs without a clear framework for converting those numbers into meaningful biological insights. This article provides a step-by-step framework for contextualizing docking scores, including units, ranges, and limitations, with practical examples drawn from published docking studies.

Docking scores are computational estimates, not measured binding affinities. They represent approximations of the free energy change associated with ligand-receptor binding, calculated through scoring functions that balance various physical and empirical terms. A more negative score generally suggests stronger predicted binding, but the absolute value carries meaning only within the context of the specific docking software, scoring function, protein preparation, and ligand library used. Without this context, comparing scores across different studies or platforms can lead to incorrect conclusions about compound potency or selectivity.

The practical outcome of this guide is a reproducible workflow for interpreting docking results: understanding what the numbers represent, establishing thresholds appropriate for your system, validating predictions through controls and orthogonal methods, and reporting results with sufficient transparency for others to evaluate. This framework applies to researchers using docking for structure-based virtual screening, lead optimization, mechanistic studies, or hypothesis generation in drug discovery projects.

## At a Glance: Docking Score Interpretation Framework

| Interpretation Step | Key Question | Practical Action |
|---------------------|--------------|------------------|
| Score normalization | What scoring function generated this number? | Record software name, version, and scoring function before comparing any values |
| Threshold establishment | What score separates likely binders from decoys in my system? | Run positive and negative controls through the same docking protocol |
| Pose evaluation | Does the top-ranked pose make chemical and biological sense? | Visually inspect hydrogen bonds, hydrophobic contacts, and steric clashes |
| Relative comparison | How does this compound compare to known actives in the same screen? | Rank compounds within a single docking run instead of across different runs |
| Validation | Does the prediction hold up in an orthogonal assay or alternative method? | Cross-check with molecular dynamics, experimental binding data, or literature precedents |
| Reporting | Can another researcher reproduce my interpretation? | Document all parameters, files, and version numbers in methods sections |

The table above summarizes the six essential steps for converting raw docking scores into defensible biological interpretations. Each step is elaborated in the sections that follow, with specific attention to the decisions researchers must make and the records they should keep.

## What Docking Scores Actually Measure

Docking scores are outputs of scoring functions that estimate the binding free energy of a ligand-protein complex. These functions approximate the change in free energy when a ligand moves from an unbound state in solution to a bound state within the protein's binding site. The approximation includes terms for van der Waals interactions, electrostatic interactions, hydrogen bonding, desolvation penalties, and entropic contributions, though the exact formulation varies substantially between different scoring functions.

The relationship between docking scores and experimentally measured binding affinities is imperfect. Scoring functions are designed to rank order compounds within a given docking experiment, not to predict absolute binding free energies with thermodynamic accuracy. A review of docking techniques in pharmacology notes that many compounds show high dock scores yet fail in preclinical studies, highlighting the gap between computational prediction and biological reality. This limitation does not diminish the utility of docking for screening and hypothesis generation, but it does require researchers to treat scores as relative indicators instead of absolute measurements.

Docking scores are expressed in arbitrary energy units, typically kilocalories per mole or related units, depending on the software. The numerical range of scores varies widely between programs. A score that indicates strong binding in one software package may be mediocre in another. This variability means that raw score values cannot be compared across different docking platforms without careful normalization or calibration.

The scoring function also depends on the docking algorithm's search strategy. Docking software typically separates the search for ligand conformations and orientations from the scoring of those poses. The search algorithm explores the conformational space of the ligand within the binding site, while the scoring function evaluates each generated pose. Different search strategies may sample different regions of conformational space, potentially missing the global energy minimum even when the scoring function would rank it favorably.

## The Role of Docking in Drug Discovery Workflows

Molecular docking occupies a specific niche in the drug discovery pipeline. It serves as a computational filter that prioritizes compounds for experimental testing, reducing the number of candidates that must be evaluated through costly and time-consuming laboratory assays. The technique has gained popularity because it saves time and money in the drug development process, particularly in the early stages where large chemical libraries must be narrowed to manageable numbers of lead candidates.

Docking is used across several distinct applications in drug discovery. Structure-based virtual screening applies docking to large compound libraries to identify potential binders for a target of interest. Hit-to-lead optimization uses docking to guide medicinal chemistry efforts by predicting how structural modifications to a lead compound might affect binding. Mechanistic studies employ docking to explore how known ligands interact with their receptors, generating hypotheses about binding modes and key interactions.

The integration of docking with other computational and experimental methods strengthens its utility. A 2025 review of molecular docking in drug discovery emphasizes that the technique's flexibility allows for the incorporation of advanced computational approaches, enhancing the reliability and efficiency of drug discovery processes. Docking results are often combined with molecular dynamics simulations, which provide more detailed information about the stability of protein-ligand complexes over time, and with experimental techniques such as surface plasmon resonance or isothermal titration calorimetry, which measure binding affinities directly.

Machine learning has begun to transform scoring function development. Recent advances include descriptor-based models and deep learning approaches that learn binding patterns from large datasets of known protein-ligand complexes. These machine learning scoring functions can sometimes outperform classical physics-based scoring functions in benchmark tests, though they bring their own challenges related to training data bias and interpretability. The choice between classical and machine learning scoring functions depends on the specific application, the availability of relevant training data, and the need for interpretability in the results.

## Units and Ranges: What the Numbers Mean

Docking scores are reported in energy units, but the specific unit depends on the software and scoring function. Common units include kilocalories per mole, kilojoules per mole, and arbitrary dimensionless units used by some empirical scoring functions. Before interpreting any docking score, researchers must confirm the units used by their software and record them in their analysis notes.

The numerical range of docking scores varies substantially across software packages. Some scoring functions produce scores in a range from approximately negative 10 to negative 2 for typical drug-like molecules, while others produce scores spanning a much wider range. The range also depends on the size and chemical properties of the ligands being docked. Larger ligands tend to produce more negative scores because they make more contacts with the protein surface, even if their binding affinity per unit of surface area is modest.

A common error in docking interpretation is treating a specific score threshold as universally meaningful. For example, a researcher might conclude that any compound with a score below negative 8 is a strong binder based on a threshold reported in a published study. This approach fails because the threshold depends on the scoring function, the protein target, the ligand library composition, and the docking protocol. A threshold that works for one system may be entirely inappropriate for another.

The appropriate way to establish score thresholds is through calibration with known actives and known inactives. If a researcher has a set of compounds with experimentally confirmed binding activity against the target, docking those compounds with the same protocol used for the screening library provides a reference distribution of scores. Compounds in the screening library that score comparably to known actives are more likely to be genuine binders than compounds that score in the range of known inactives.

## Protein and Ligand Preparation: The Foundation of Meaningful Scores

Docking scores are only as reliable as the structures fed into the docking calculation. Protein preparation involves obtaining a high-quality three-dimensional structure of the target, typically from experimental methods such as X-ray crystallography, cryo-electron microscopy, or nuclear magnetic resonance spectroscopy. The structure must be checked for completeness, with missing loops or side chains either resolved or noted as limitations.

The choice of binding site is a critical decision that directly affects docking scores. Docking a ligand into the wrong site produces scores that are meaningless for the intended biological question. Binding sites can be identified from co-crystallized ligands, from computational pocket detection algorithms, or from mutagenesis studies that identify residues essential for ligand binding. Each approach has limitations, and the selected site should be justified based on the available evidence.

Ligand preparation includes generating three-dimensional coordinates, assigning protonation states at physiological pH, and generating stereoisomers and tautomers. The protonation state of ionizable groups significantly affects electrostatic interactions and hydrogen bonding patterns, which in turn affect docking scores. Many ligands contain functional groups that can exist in multiple protonation states, and the correct state depends on the local environment within the binding site.

Water molecules in the binding site present a particular challenge. Some crystallographic water molecules mediate protein-ligand interactions and should be retained during docking, while others are displaced upon ligand binding and should be removed. The decision to retain or remove specific water molecules can change docking scores and ranking. Researchers should document their water handling strategy and consider testing multiple configurations to assess the sensitivity of their results.

## Selecting the Right Docking Software and Scoring Function

The choice of docking software and scoring function is one of the most consequential decisions in any docking study. Different programs implement different search algorithms and scoring functions, and their performance varies depending on the protein target and ligand properties. No single program performs best across all systems, and published benchmarks often show that performance is system-dependent.

Classical scoring functions are based on physical principles, including molecular mechanics force fields, empirical terms derived from fitting to experimental binding data, and knowledge-based potentials derived from statistical analysis of known protein-ligand complexes. Each approach has strengths and weaknesses. Force field-based functions capture detailed atomic interactions but are computationally expensive. Empirical functions are fast and often perform well for their training set but may not generalize to novel chemotypes. Knowledge-based potentials are computationally efficient but depend heavily on the quality and diversity of the structural data used to derive them.

Machine learning scoring functions represent a newer category that has gained substantial attention. These functions are trained on large datasets of protein-ligand complexes with known binding affinities or binding classifications. Deep learning approaches can automatically extract relevant features from atomic coordinates, potentially capturing patterns that are difficult to encode in classical scoring functions. However, machine learning scoring functions are only as good as their training data, and they may perform poorly on protein targets or ligand chemotypes that are underrepresented in the training set.

The choice between software packages also involves practical considerations. Some programs are freely available for academic use, while others require commercial licenses. Computational cost varies substantially, with some programs capable of docking thousands of compounds per hour on standard hardware and others requiring significantly more time. The availability of documentation, user communities, and technical support also influences the practical usability of different programs.

## Establishing Controls and Calibration Sets

Controls are essential for meaningful docking score interpretation. A docking experiment without appropriate controls produces numbers that cannot be reliably interpreted. The most important controls are positive controls, which are compounds with known binding activity against the target, and negative controls, which are compounds known not to bind or which are structurally similar decoys.

Positive controls serve multiple purposes in docking studies. They validate that the docking protocol can reproduce known binding poses when the co-crystallized ligand is re-docked into its binding site. They establish the score range expected for genuine binders. They also provide a reference point for ranking screening hits. If a known active with a measured binding affinity in the micromolar range produces a docking score of negative 8, then screening compounds with scores of negative 9 or lower are likely to have affinities in a similar or better range.

Negative controls help establish the background distribution of docking scores for compounds that do not bind. Decoy sets, which are compounds with similar physical properties to known actives but different chemical structures, are commonly used for this purpose. The separation between the score distributions of known actives and decoys provides a measure of the docking protocol's ability to discriminate between binders and non-binders. Poor separation indicates that the protocol may not be suitable for virtual screening against this target.

Redocking experiments, where the crystallographic ligand is removed from the binding site and docked back into the structure, assess the docking protocol's ability to reproduce the experimentally observed binding pose. The root mean square deviation between the docked pose and the crystallographic pose provides a quantitative measure of pose prediction accuracy. A root mean square deviation below a certain threshold, typically around 2 angstroms, is often used as a criterion for successful pose reproduction, though the appropriate threshold depends on the specific system and research question.

## Practical Workflow for Docking Score Interpretation

A reproducible workflow for docking score interpretation begins before any docking calculation is performed. The first step is to define the biological question and the decision that the docking results will inform. Is the goal to select compounds for experimental testing, to prioritize analogs for synthesis, or to generate hypotheses about binding modes? The answer determines the appropriate level of rigor and the types of validation required.

The second step is to assemble the necessary input data. This includes the protein structure, the ligand library, and any experimental data that can serve as controls. The protein structure should be checked for quality, including resolution, completeness, and the presence of any ligands or cofactors that might affect the binding site. The ligand library should be prepared with consistent protonation states, stereochemistry, and tautomeric forms.

The third step is to run the docking calculations with documented parameters. All settings should be recorded, including the search algorithm, the number of docking runs per ligand, the exhaustiveness or sampling parameters, and the scoring function. These parameters should be consistent across all ligands in the study to ensure that scores are comparable.

The fourth step is to analyze the results with attention to both scores and poses. Compounds should be ranked by score, but the top-ranked poses should also be visually inspected for reasonable geometry, complementary interactions, and the absence of severe steric clashes. A compound with an excellent score but an implausible pose should be treated with suspicion.

The fifth step is to validate the interpretation through orthogonal methods. This may include molecular dynamics simulations to assess the stability of the predicted binding mode, comparison with known structure-activity relationships, or experimental binding assays for the most promising compounds. The level of validation should match the significance of the decision being made.

## Records and Measurements: Documenting Docking Interpretations

Systematic record keeping is essential for reproducible docking studies. The records should capture all information needed for another researcher to reproduce the docking calculations and the interpretation of the results. This includes software versions, parameter settings, input file preparation steps, and the rationale for key decisions.

A docking study record should include the following elements: the protein structure identifier and source, the preparation steps applied to the protein, the binding site definition and its justification, the ligand library source and preparation protocol, the docking software name and version, the scoring function and all adjustable parameters, the number of docking runs and the selection criteria for the final pose, and the control compounds and their results.

The interpretation of docking scores should also be documented. This includes the score distributions for known actives and decoys, the threshold selected for classifying potential hits, and the rationale for that threshold. If the threshold is based on published precedents, those precedents should be cited. If the threshold is based on the control compounds in the current study, the control results should be reported.

The limitations of the docking study should be recorded alongside the results. This includes any uncertainties in the protein structure, such as missing loops or low resolution, any assumptions about protonation states or tautomers, and any known limitations of the scoring function for the specific chemical space being explored. Transparent reporting of limitations allows other researchers to assess the reliability of the conclusions.

## Common Failure Patterns in Docking Score Interpretation

Several recurring errors undermine the validity of docking score interpretations. Recognizing these failure patterns helps researchers avoid them and helps reviewers identify problematic studies.

The first common failure is comparing docking scores across different software packages or scoring functions. A score of negative 9 from one program does not have the same meaning as a score of negative 9 from another program. Even within the same software, changing the scoring function or the protein preparation protocol can shift the score distribution substantially. Scores should only be compared within a single docking run using consistent protocols.

The second failure is treating docking scores as quantitative predictions of binding affinity. Docking scores are rank-ordering tools, not thermodynamic measurements. A compound with a docking score of negative 10 is not necessarily a stronger binder than a compound with a score of negative 9. The difference between scores does not correspond to a specific difference in binding free energy. Quantitative affinity predictions require more rigorous methods such as free energy perturbation calculations or experimental measurements.

The third failure is ignoring the pose when interpreting the score. A docking score is calculated for a specific pose, and the score is only meaningful if that pose is physically reasonable. A pose with severe steric clashes, unsatisfied hydrogen bond donors or acceptors buried in hydrophobic regions, or strained bond geometries should not be accepted even if the score is favorable. Visual inspection of poses is an essential quality control step.

The fourth failure is overinterpreting results from a single docking run. Docking algorithms are stochastic, meaning that repeated runs may produce different poses and scores. Running multiple docking calculations for each ligand and examining the consistency of the results provides a measure of confidence. Ligands that produce consistent poses across multiple runs are more reliable than those that produce widely varying poses.

The fifth failure is neglecting to validate the docking protocol for the specific target. A protocol that works well for one protein may perform poorly for another. The only way to know whether a docking protocol is appropriate for a given target is to test it with known actives and decoys. Skipping this validation step produces results of unknown reliability.

## Limitations of Docking Scores and Their Interpretation

Docking scores have inherent limitations that cannot be fully overcome through careful protocol design. Understanding these limitations is essential for interpreting results appropriately and for communicating the uncertainty to collaborators and decision makers.

The treatment of protein flexibility is a major limitation. Most docking protocols treat the protein as rigid or allow only limited flexibility of selected side chains. In reality, proteins undergo conformational changes upon ligand binding, and these changes can significantly affect binding affinity. Induced fit effects, where the protein adapts its shape to accommodate the ligand, are poorly captured by most docking protocols. This limitation is particularly important for targets with flexible binding sites or for ligands that bind through conformational selection mechanisms.

The treatment of solvation and desolvation is another significant limitation. Scoring functions use simplified models of water, often implicit solvation models that approximate the effects of water as a continuous medium. The displacement of ordered water molecules from the binding site upon ligand binding contributes to the binding free energy, but this effect is difficult to model accurately. Ligands that make favorable interactions with water in the unbound state may pay a larger desolvation penalty upon binding, and this penalty may be underestimated by the scoring function.

Entropic effects are also poorly captured by docking scores. The loss of conformational freedom when a ligand binds to a protein contributes unfavorably to the binding free energy. Flexible ligands with many rotatable bonds pay a larger entropic penalty than rigid ligands. Most scoring functions either ignore this term entirely or approximate it with simple corrections based on the number of rotatable bonds. The result is that docking scores may overestimate the binding affinity of flexible ligands.

The accuracy of the input protein structure imposes a fundamental limit on docking accuracy. Docking into a structure with errors, such as incorrectly assigned side chain conformations or inaccurate loop regions, produces unreliable results regardless of the quality of the scoring function. The resolution of the experimental structure, the presence of crystal contacts that may distort the binding site, and the relevance of the crystallographic conformation to the solution state all affect the reliability of docking results.

## Machine Learning Scoring Functions: New Capabilities and New Caveats

The integration of machine learning into docking scoring functions represents a significant development in the field. Machine learning scoring functions are trained on datasets of protein-ligand complexes with known binding data, learning patterns that relate structural features to binding outcomes. These functions can capture complex relationships that are difficult to encode in classical physics-based scoring functions.

Descriptor-based machine learning scoring functions use precomputed features of the protein-ligand complex, such as counts of specific interaction types, surface area descriptors, or physicochemical properties. These features are fed into machine learning models such as random forests, support vector machines, or gradient boosting machines. The models learn to map the descriptor values to binding affinity or binding classification.

Deep learning scoring functions take a different approach, operating directly on atomic coordinates or grid representations of the protein-ligand complex. Convolutional neural networks can learn spatial patterns of interactions, while graph neural networks can represent the molecular structure as a graph and learn to predict binding from the graph topology. These approaches can potentially capture features that are not explicitly encoded in descriptor-based methods.

The performance of machine learning scoring functions depends critically on the training data. Datasets such as the Protein Data Bank bind database provide large collections of protein-ligand complexes with measured binding affinities, but these datasets have biases. Certain protein families, ligand chemotypes, and affinity ranges are overrepresented, and the models may perform poorly on targets or compounds that differ from the training distribution. The risk of overfitting to training data is a particular concern for deep learning approaches with large numbers of parameters.

The interpretability of machine learning scoring functions is another consideration. Classical scoring functions have transparent functional forms, allowing researchers to understand why a particular compound receives a particular score. Machine learning models, particularly deep learning models, are often opaque, making it difficult to diagnose why a prediction is made. This lack of interpretability can complicate the use of docking results in medicinal chemistry decision making, where understanding the structural basis of predicted binding is valuable.

## Case Example: Docking in Autoimmune Disease Research

A 2026 study of thyroid-stimulating hormone receptor in autoimmune thyroid diseases illustrates how docking scores are used in contemporary biomedical research. The study integrated genome-wide association studies, transcriptomics, and brain imaging phenotypes to characterize peripheral and central alterations in Graves' disease and Graves' orbitopathy. Structural biology analyses included protein-protein docking, small-molecule docking, and normal mode dynamics to identify prospective modulators of the thyroid-stimulating hormone receptor.

In this context, docking scores served as one component of a multi-layered investigation. The docking calculations were used to identify potential modulators of the receptor, but these predictions were not treated as standalone evidence. The study also employed Mendelian randomization to test causal relationships between genetic variants and brain signatures, and immunofluorescence staining in a mouse model to validate the colocalization of potential interacting proteins in specific brain regions.

This example demonstrates the appropriate use of docking scores in modern research: as hypothesis-generating tools that are validated through orthogonal methods. The docking predictions identified candidate modulators, but the biological significance of those candidates was established through genetic, transcriptomic, and experimental validation. Researchers who treat docking scores as definitive evidence of binding or biological activity are misusing the technique.

A second example comes from a 2025 study of Gui-Zhi-Shao-Yao-Zhi-Mu decoction in rheumatoid arthritis. The study used high-throughput sequencing data, Chinese Materia Medica target-related databases, and ultra-high-performance liquid chromatography coupled with high-resolution mass spectrometry to identify active compounds. The research constructed a Chinese Materia Medica-Ingredient-Target network to obtain candidate drug target genes and identified hub genes based on a Molecular Complex Detection algorithm.

In this study, docking would typically be used to explore how the identified compounds interact with the candidate target proteins. The docking scores would help prioritize which compound-target pairs to investigate experimentally. However, the biological conclusions of the study would rest on the experimental validation, not on the docking scores alone. This integration of computational prediction with experimental confirmation represents the standard of practice in the field.

## Reporting Docking Results for Reproducibility

The reporting of docking results in publications and reports should follow standards that enable other researchers to evaluate and reproduce the work. The level of detail required depends on the purpose of the report, but several elements are essential for any docking study that informs scientific conclusions.

The protein structure should be identified by its database accession code and the specific chain or conformation used. The preparation steps applied to the protein should be described, including any addition of hydrogen atoms, assignment of protonation states, and treatment of water molecules. The binding site should be defined, either by the coordinates of the region used for docking or by the residues that define the site.

The ligand preparation protocol should be described, including the source of the ligand structures, the generation of three-dimensional coordinates, and the treatment of stereochemistry and tautomers. The docking software and version should be identified, along with the scoring function and all adjustable parameters. The number of docking runs per ligand and the criteria for selecting the final pose should be stated.

The results should include the score distributions for the compounds studied, beyond the top-scoring hits. Reporting the full distribution allows readers to assess the separation between potential hits and the background. The results for control compounds should be reported, including the scores for known actives and decoys. The threshold used to classify potential hits should be stated with its justification.

The limitations of the study should be discussed explicitly. This includes any known issues with the protein structure, any assumptions made during preparation, and any concerns about the applicability of the scoring function to the chemical space studied. A study that acknowledges its limitations is more credible than one that presents docking scores as definitive measurements.

## Professional Escalation Criteria for Docking Interpretation

Researchers should recognize when docking score interpretation requires consultation with specialists or additional investigation. Several situations warrant escalation to colleagues with deeper expertise in computational chemistry, structural biology, or the specific target system.

When docking results are inconsistent with experimental data, escalation is appropriate. If a compound with a favorable docking score fails to show binding in experimental assays, or if a compound with an unfavorable score shows strong binding, the discrepancy should be investigated instead of ignored. Possible explanations include errors in the docking protocol, incorrect assumptions about the binding site, or limitations of the scoring function for the specific system.

When the protein structure has significant uncertainties, consultation with a structural biologist is advisable. This includes structures with low resolution, missing regions, or conformational heterogeneity. A structural biologist can help assess whether the structure is suitable for docking and whether alternative conformations should be considered.

When the docking results will inform major decisions, such as the selection of compounds for expensive experimental studies or the progression of a drug discovery program, the interpretation should be reviewed by multiple experts. The cost of a false positive or false negative at this stage is high, and the additional perspective provided by expert review can identify issues that a single researcher might miss.

When machine learning scoring functions are used, the limitations of the training data should be carefully assessed. If the target protein or the ligand chemotypes being studied are poorly represented in the training data, the predictions may be unreliable. Consultation with a machine learning specialist can help assess the applicability domain of the model and the confidence that can be placed in its predictions.

## Frequently Asked Questions

### What is a good docking score?

A good docking score is one that falls within the range established for known active compounds against your specific target using your specific docking protocol. There is no universal threshold that separates good from poor scores across all systems. The appropriate way to determine what constitutes a good score is to dock known actives and known inactives with the same protocol used for your screening library, then examine the separation between the two score distributions. A screening compound that scores in the range of known actives is a candidate for experimental testing.

### Can docking scores be compared across different software programs?

Docking scores should not be compared across different software programs or scoring functions. Each program uses a different scoring function with different units, ranges, and parameterizations. A score that indicates strong binding in one program may be mediocre in another. Even within the same program, changing the protein preparation protocol or the scoring function parameters can shift the score distribution. Scores should only be compared within a single docking run using consistent protocols.

### How accurate are docking scores for predicting binding affinity?

Docking scores are rank-ordering tools instead of quantitative predictions of binding affinity. They are useful for prioritizing compounds within a screening library, but they do not provide accurate estimates of absolute binding free energies. The correlation between docking scores and experimentally measured binding affinities is typically modest, and the relationship varies substantially across different protein targets and ligand chemotypes. Quantitative affinity predictions require more rigorous computational methods or experimental measurements.

### Why do compounds with good docking scores fail in experimental assays?

Compounds with favorable docking scores can fail in experimental assays for many reasons. The docking score may be an artifact of the scoring function, which may not accurately capture the true binding free energy. The predicted binding pose may not be the pose adopted in solution. The compound may have poor solubility, permeability, or metabolic stability that prevents it from reaching the target in the assay. The docking calculation may have used a protein conformation that is not relevant under physiological conditions. These possibilities should be investigated when docking predictions do not match experimental results.

### How many docking runs should be performed for each ligand?

The number of docking runs per ligand depends on the software and the variability of the results. Running multiple docking calculations for each ligand and examining the consistency of the poses and scores provides a measure of confidence. If repeated runs produce consistent poses with similar scores, fewer runs may be sufficient. If the results vary substantially between runs, more runs are needed to characterize the distribution of possible outcomes. The number of runs should be documented in the methods.

### What is the role of molecular dynamics simulations in validating docking results?

Molecular dynamics simulations can assess the stability of a predicted protein-ligand binding mode over time. A docking pose that is stable during a molecular dynamics simulation, maintaining key interactions and remaining within the binding site, is more credible than a pose that quickly dissociates or undergoes major conformational changes. Molecular dynamics can also reveal induced fit effects and alternative binding modes that are not captured by rigid receptor docking. However, molecular dynamics simulations are computationally expensive and require careful setup and analysis.

### How should water molecules be treated in docking calculations?

The treatment of water molecules in docking depends on the specific binding site and the role of water in ligand binding. Crystallographic water molecules that mediate protein-ligand interactions should generally be retained, while water molecules that would be displaced upon ligand binding should be removed. The decision should be based on the structural evidence and should be documented. Testing multiple water configurations can assess the sensitivity of the docking results to water handling.

### When should machine learning scoring functions be used instead of classical scoring functions?

Machine learning scoring functions may offer improved predictive performance when the target protein and ligand chemotypes are well represented in the training data. They may be particularly useful for virtual screening applications where ranking accuracy is more important than interpretability. However, classical scoring functions may be preferred when interpretability is important, when the target is unusual or poorly represented in training data, or when computational resources are limited. The choice should be based on benchmarking against known actives and decoys for the specific system of interest.

## Related Bioinformatics Guides

- [Metagenome Assembled Genome Analysis: From Bins to Biological Insights](/knowledge/bioinformatics/metagenome-assembled-genome-analysis-from-bins-to-biological-insights)
- [Genomic Data Integration: Combining Multi-Omics for Biological Insights](/knowledge/bioinformatics/genomic-data-integration-combining-multi-omics-for-biological-insights)
- [Metagenomics Data Analysis: From Raw Reads to Biological Insights](/knowledge/bioinformatics/metagenomics-data-analysis-from-raw-reads-to-biological-insights)
- [Proteomics Data Analysis Workflow: From Raw Spectra to Biological Insights](/knowledge/bioinformatics/proteomics-data-analysis-workflow-from-raw-spectra-to-biological-insights)
- [Genomic Data Analytics: Extracting Biological Insights from Large-Scale Sequencing](/knowledge/bioinformatics/genomic-data-analytics-extracting-biological-insights-from-large-scale-sequencing)

## References and Further Reading

- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Bioconductor](https://bioconductor.org/). Bioconductor Project.
- [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.
- [nf-core Documentation](https://nf-co.re/docs). nf-core.
- [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.
- [Molecular Docking in Drug Discovery: Techniques, Applications, and Advancements.](https://pubmed.ncbi.nlm.nih.gov/39415575). Current medicinal chemistry, 2025.
- [Bioinformatics identification based on causal association inference using multi-omics reveals the underlying mechanism of Gui-Zhi-Shao-Yao-Zhi-Mu decoction in modulating rheumatoid arthritis.](https://pubmed.ncbi.nlm.nih.gov/39736250). Phytomedicine : international journal of phytotherapy and phytopharmacology, 2025.
- [Docking techniques in pharmacology: How much promising?](https://pubmed.ncbi.nlm.nih.gov/30067954). Computational biology and chemistry, 2018.
- [Thyroid-stimulating hormone receptor mediates peripheral-central neuroimmune crosstalk in autoimmune thyroid diseases.](https://pubmed.ncbi.nlm.nih.gov/42204533). BMC medicine, 2026.
- [Protein-Ligand Docking in the Machine-Learning Era.](https://pubmed.ncbi.nlm.nih.gov/35889440). Molecules (Basel, Switzerland), 2022.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.