# How to Interpret DALI Z-Scores: What They Really Tell You About Structural Similarity

## Direct Answer and Scope

DALI Z-scores quantify the statistical significance of protein structural similarity by expressing how many standard deviations a pairwise structural alignment score sits above the mean score expected from random protein pairs of comparable size and shape. A Z-score of 2.0 indicates that the observed structural similarity is approximately two standard deviations above what random chance would produce, while scores above 8.0 typically represent confident structural matches that merit detailed biological investigation. This article explains the statistical foundation of DALI Z-scores, how they relate to structural similarity and significance, and provides practical thresholds and examples for interpreting results in research and laboratory settings.

The intended readers are biology students, researchers, laboratory professionals, and life-science practitioners who receive DALI output files and need to make defensible decisions about whether detected structural similarities represent genuine biological relationships or statistical artifacts. The scope covers the mathematical meaning of Z-scores, the relationship between Z-scores and other structural similarity metrics, practical interpretation thresholds, common interpretation errors, and reproducible workflow practices for structural bioinformatics analysis.

## The Statistical Foundation of DALI Z-Scores

### What a Z-Score Measures in Structural Alignment

A Z-score in the context of DALI structural alignment represents the number of standard deviations by which an observed structural similarity score exceeds the mean score of a reference distribution of random structural comparisons. The reference distribution is constructed from many pairwise alignments between structurally unrelated proteins, and the observed score for a query pair is positioned within that distribution to determine how unusual or significant the match is.

The statistical logic follows the standard normal distribution framework. If random protein pairs produce alignment scores that cluster around a mean with a measurable spread, then a query pair whose score falls far into the upper tail of that distribution is unlikely to have arisen by chance. The Z-score converts this position into a standardized value that can be compared across different protein pairs, different database searches, and different alignment parameters.

The key distinction from raw alignment scores is that Z-scores account for background expectations. A raw score that seems high in absolute terms may be unremarkable if random protein pairs frequently achieve similar scores. Conversely, a modest raw score may be highly significant if the reference distribution is narrow and the query pair sits far in the upper tail. This normalization is what makes Z-scores useful for comparing significance across alignments of proteins with different lengths, shapes, and structural classes.

### The Role of Random Models in Significance Assessment

Quantification of statistical significance is essential for the interpretation of protein structural similarity, and random models for protein structure comparison provide the foundation for this quantification. A random model for protein structure comparison was developed with three distinctive features. First, a sample of random structure comparisons is restricted to molecules of the same size and shape as the superposition of interest. Second, careful selection of the sample and accurate modeling of shape allows approximation of the root mean square deviation distribution of random comparisons with a Nakagami probability density function. Third, through convolution, a second probability density function is obtained that describes the coordinate difference vector projections underlying the random distribution of root mean square deviation.

This modeling approach allows sampling of random distributions for any similarity score that depends on difference vector projections, including the GDT_TS score, the TM score, and the LiveBench 3D score. Probabilities estimated from this method correlate well with common measures of structural similarity, such as the DALI Z-score and the GDT_TS score. The practical consequence is that a p-value for a given superposition can be calculated using simple formulae depending on root mean square deviation, radius of gyration, and thinnest molecular dimension.

The connection between Z-scores and p-values is direct. A Z-score of 2.0 corresponds approximately to a p-value of 0.05 in a one-tailed normal distribution, meaning that about 5 percent of random comparisons would produce a score this high or higher. A Z-score of 3.0 corresponds to a p-value near 0.001, and a Z-score of 8.0 corresponds to an extremely small p-value that indicates the observed similarity is essentially impossible under the random model.

### Why Size and Shape Matter in Z-Score Calculation

The random model approach emphasizes that structural comparison statistics must account for the size and shape of the molecules being compared. A large protein domain has more opportunities to achieve a given root mean square deviation by chance than a small domain, simply because there are more atoms and more possible superpositions. Similarly, elongated proteins have different random alignment properties than globular proteins because their shape constrains the possible superpositions.

The DALI Z-score inherits this size and shape dependence through its underlying score distribution. When interpreting a Z-score, researchers should consider whether the proteins being compared are similar in size and shape to the proteins used to construct the reference distribution. Comparisons between proteins of very different sizes may produce Z-scores that are not directly comparable to Z-scores from same-size comparisons.

This size dependence has practical consequences for database searches. A query protein that is unusually large or small relative to the database population may produce Z-scores that are systematically inflated or deflated. Researchers should examine the distribution of Z-scores across all hits in a search, beyond the top hit, to calibrate their expectations for their specific query protein.

## DALI Z-Scores in the Context of Structural Similarity Metrics

### Relationship Between Z-Scores and Root Mean Square Deviation

Root mean square deviation measures the average distance between corresponding atoms after optimal superposition of two structures. Lower root mean square deviation values indicate greater structural similarity. However, root mean square deviation alone is insufficient for assessing significance because it does not account for the number of aligned residues or the background distribution of random alignments.

The random model approach addresses this limitation by providing a statistical framework for root mean square deviation interpretation. The root mean square deviation distribution of random comparisons can be approximated with a Nakagami probability density function, and this distribution depends on the size and shape of the molecules being compared. A root mean square deviation of 2 angstroms between two 50-residue domains may be highly significant, while the same root mean square deviation between two 500-residue domains may be less remarkable because larger proteins have more opportunities to achieve low root mean square deviation by chance.

DALI Z-scores integrate root mean square deviation information with alignment length and structural continuity information into a single significance measure. When interpreting a DALI result, researchers should examine both the Z-score and the root mean square deviation to understand the nature of the structural similarity. A high Z-score with a low root mean square deviation indicates a confident, tight structural match. A high Z-score with a moderate root mean square deviation may indicate a confident match over a large structural region with some local deviations.

### Relationship Between Z-Scores and GDT_TS and TM Scores

The GDT_TS score measures the percentage of residues that can be superimposed within specific distance thresholds, typically 1, 2, 4, and 8 angstroms. The TM score normalizes structural similarity by protein length and provides a length-independent measure of topological similarity. Both metrics are widely used in structure prediction assessment and structural comparison.

The random model approach demonstrates that probabilities estimated from the method correlate well with common measures of structural similarity, including the DALI Z-score and the GDT_TS score. This correlation means that these different metrics are capturing related aspects of structural similarity, but they emphasize different properties. The DALI Z-score emphasizes statistical significance relative to random expectation, while GDT_TS and TM scores emphasize the geometric quality of the superposition.

For practical interpretation, researchers should use Z-scores to assess whether a structural similarity is statistically meaningful and use GDT_TS or TM scores to assess the geometric quality of the alignment. A high Z-score with a low GDT_TS score may indicate that the structural similarity is statistically significant but geometrically imperfect, perhaps because the similarity is confined to a substructure or because the proteins have different overall folds with a conserved core.

### The Complementary Role of P-Values

P-values provide a direct probability interpretation of structural similarity significance. The random model approach enables calculation of p-values for a given superposition using simple formulae depending on root mean square deviation, radius of gyration, and thinnest molecular dimension. These p-values offer a statistically sound alternative to scores used in reference-independent evaluation of alignment quality.

The relationship between Z-scores and p-values is monotonic but not linear. A Z-score of 2.0 corresponds to a p-value of approximately 0.05, a Z-score of 3.0 corresponds to a p-value of approximately 0.001, and higher Z-scores correspond to exponentially smaller p-values. Researchers who need to report significance in terms of p-values, for example in publications or regulatory submissions, can convert Z-scores to approximate p-values using standard normal distribution tables or statistical software.

The practical advantage of p-values is that they have a direct probabilistic interpretation that is familiar to a broad scientific audience. The practical advantage of Z-scores is that they are the native output of DALI and provide a continuous measure that preserves information about the degree of significance beyond simple threshold categories.

## Practical Thresholds for DALI Z-Score Interpretation

### Threshold Categories for Research Decisions

DALI Z-scores are commonly interpreted using threshold categories that guide research decisions. A Z-score below 2.0 generally indicates that the structural similarity is not statistically significant and could easily arise by chance. A Z-score between 2.0 and 4.0 indicates weak but potentially interesting similarity that warrants further investigation. A Z-score between 4.0 and 8.0 indicates confident structural similarity that likely reflects a genuine evolutionary or functional relationship. A Z-score above 8.0 indicates highly confident structural similarity that is essentially impossible under the random model.

These thresholds are practical guidelines instead of rigid rules. The appropriate threshold for a given research question depends on the context. A researcher screening a large database for potential homologs may use a higher threshold to reduce false positives, while a researcher investigating a specific pair of proteins with prior biological evidence for similarity may accept a lower threshold.

The threshold categories should be applied with attention to the size and shape of the proteins being compared. The random model approach emphasizes that significance assessment must account for molecular size and shape, and the same Z-score may have different practical implications for different protein classes.

### At a Glance: DALI Z-Score Interpretation Table

| Z-Score Range | Significance Level | Practical Interpretation | Recommended Action |
| --- | --- | --- | --- |
| Below 2.0 | Not significant | Similarity is within random expectation | Report as no significant structural similarity |
| 2.0 to 4.0 | Weak | Possible local or partial similarity | Examine alignment details and biological context |
| 4.0 to 8.0 | Confident | Likely genuine structural relationship | Proceed with detailed structural analysis |
| Above 8.0 | Highly confident | Essentially certain structural relationship | Prioritize for functional and evolutionary investigation |

### Context-Dependent Threshold Adjustments

The standard threshold categories require adjustment for specific research contexts. In large-scale database searches where thousands of comparisons are performed, the multiple testing problem means that some high Z-scores will arise by chance. A Z-score of 4.0 that is highly significant in a single pairwise comparison may be less compelling when it appears among thousands of database hits.

Researchers performing DALI searches against large databases should examine the distribution of Z-scores across all hits and compare the top hits to the background distribution. If many hits have Z-scores above 4.0, the effective threshold for confident significance should be raised. If the query protein is structurally unusual, the background distribution may be shifted, and the standard thresholds may not apply.

The size and shape dependence of the random model also affects threshold interpretation. The random model restricts comparisons to molecules of the same size and shape as the superposition of interest, and this restriction affects the expected distribution of scores. Researchers comparing proteins of very different sizes should be cautious about applying standard thresholds and should consider generating a size-matched background distribution for their specific comparison.

## The DALI Search Workflow and Data Inputs

### Preparing Structural Data for DALI Analysis

DALI analysis requires protein structure data in a format that the DALI server or standalone software can process. The most common input format is the Protein Data Bank format, which contains atomic coordinates for the protein structure. Researchers can obtain structure files from the Protein Data Bank or from their own structural determination experiments.

Before running DALI, researchers should verify the quality of their input structures. Structures with missing residues, poor electron density, or multiple conformations may produce misleading alignment results. The DALI algorithm aligns structural elements based on their three-dimensional coordinates, and errors in the input coordinates propagate into the alignment and the resulting Z-scores.

For structures determined by cryo-electron microscopy or nuclear magnetic resonance, researchers should consider which model or conformer to use for DALI analysis. Different conformers of the same protein may produce different Z-scores against the same database, and the choice of conformer should be documented in the analysis protocol.

### Running DALI Searches Against Structural Databases

The DALI server accepts a query structure and searches it against a database of known protein structures. The default database includes a comprehensive collection of protein structures from the Protein Data Bank, and the search returns a ranked list of structural neighbors with their Z-scores, root mean square deviation values, and alignment details.

Researchers should document the database version and search parameters used for each DALI search. Different database versions contain different sets of structures, and the Z-scores for a given query can change when the database is updated. The search parameters, including the minimum alignment length and the structural alphabet settings, also affect the results.

For reproducible research, the DALI search should be repeated with the same parameters and database version, and the results should be archived with the analysis protocol. The Galaxy Training Network provides accessible workflow training and analysis tutorials that emphasize reproducibility context for bioinformatics analyses, and similar principles apply to structural bioinformatics workflows.

### Recording DALI Results for Analysis

The DALI output includes a ranked list of structural neighbors with their Z-scores, root mean square deviation values, sequence identities, and alignment lengths. Researchers should record the complete output for their analysis, beyond the top hits, because the distribution of Z-scores across all hits provides important context for interpreting the significance of individual matches.

For each significant hit, researchers should record the Z-score, the root mean square deviation, the number of aligned residues, the percentage sequence identity, and the structural classification of the matched protein. This information supports downstream analysis of whether the structural similarity reflects evolutionary relatedness, functional convergence, or shared structural constraints.

The records should include the date of the search, the database version, the query structure identifier, and the software version. These metadata are essential for reproducing the analysis and for comparing results across different searches or different research projects.

## Interpreting DALI Z-Scores in Different Research Contexts

### Structure Prediction Assessment

DALI Z-scores are used to assess the quality of predicted protein structures by comparing them to experimentally determined structures. A predicted structure that achieves a high Z-score against the experimental structure indicates that the prediction captured the correct fold. The random model approach provides a statistically sound alternative to scores used in reference-independent evaluation of alignment quality, and this approach can be applied to evaluation of homology modeling techniques.

When assessing structure predictions, researchers should compare the Z-score of the prediction against the experimental structure to the Z-scores of the prediction against unrelated structures. A prediction that achieves a high Z-score only against the target structure, and low Z-scores against all other structures, has captured the correct fold. A prediction that achieves moderate Z-scores against many structures may have captured a common fold but not the specific details of the target structure.

The GDT_TS score is commonly used in structure prediction assessment, and the correlation between GDT_TS and DALI Z-scores means that both metrics provide complementary information. Researchers should report both metrics when assessing prediction quality, along with the root mean square deviation and the alignment length.

### Evolutionary and Functional Inference

High DALI Z-scores between proteins with low sequence identity provide evidence for evolutionary relatedness through structural conservation. Proteins that share a common ancestor but have diverged beyond sequence recognition often retain similar three-dimensional structures, and DALI can detect these relationships.

When using DALI Z-scores for evolutionary inference, researchers should consider the structural classification of the matched proteins. Proteins in the same structural class or fold family are more likely to share evolutionary ancestry than proteins that merely share a common structural motif. The Z-score provides evidence for structural similarity, but the evolutionary interpretation requires additional context from functional data, phylogenetic analysis, and structural classification databases.

The NCBI Data Resources provide official descriptions of databases, search systems, sequence resources, and analysis services that can support evolutionary analysis of DALI hits. Researchers can use these resources to retrieve sequence information, functional annotations, and taxonomic distributions for the proteins identified through DALI searches.

### Functional Annotation Transfer

High-confidence DALI hits can support functional annotation transfer when the structural similarity is strong and the matched protein has well-characterized function. The structural similarity suggests that the query protein may have a similar function, particularly if the functionally important residues are conserved in the alignment.

The confidence in functional annotation transfer depends on the Z-score, the alignment quality, and the functional relevance of the conserved structural features. A high Z-score with a long alignment and conserved active site residues provides strong evidence for functional similarity. A high Z-score with a short alignment or divergent functional residues provides weaker evidence.

Researchers should document the evidence chain for functional annotation transfer, including the DALI Z-score, the alignment details, the functional annotations of the matched protein, and any experimental evidence supporting the functional assignment. This documentation supports scientific rigor and enables other researchers to evaluate the strength of the functional inference.

## Common Failure Patterns in DALI Z-Score Interpretation

### Overinterpreting Marginal Z-Scores

A common failure pattern is treating a Z-score above 2.0 as evidence of genuine structural similarity without considering the multiple testing context. In a database search against thousands of structures, a Z-score of 2.0 to 3.0 may arise for several hits purely by chance. The statistical significance of an individual hit must be interpreted in the context of the total number of comparisons performed.

Researchers should examine the distribution of Z-scores across all hits in a DALI search. If the distribution shows a smooth decay with many hits in the 2.0 to 4.0 range, these marginal hits are likely part of the random background. If the distribution shows a clear gap between a few high-scoring hits and the rest of the distribution, the high-scoring hits are more likely to represent genuine structural relationships.

The random model approach provides a principled way to assess whether a given Z-score is significant in the context of the search. The p-value for a given superposition can be calculated using simple formulae depending on root mean square deviation, radius of gyration, and thinnest molecular dimension, and this p-value accounts for the specific properties of the comparison.

### Ignoring Alignment Quality Details

A second common failure pattern is focusing exclusively on the Z-score while ignoring the alignment quality details. Two alignments with the same Z-score can have very different properties. One alignment may cover 90 percent of both proteins with low root mean square deviation, while another may cover only 30 percent of both proteins with higher root mean square deviation.

The alignment length and the root mean square deviation provide essential context for interpreting the Z-score. A high Z-score with a short alignment indicates that the structural similarity is confined to a substructure, which may represent a shared domain or a conserved structural motif. A high Z-score with a long alignment indicates that the overall folds are similar, which provides stronger evidence for evolutionary relatedness.

Researchers should examine the structural superposition of significant hits to verify that the alignment makes biological sense. Automated alignments can sometimes align secondary structure elements in ways that are geometrically valid but biologically meaningless, particularly when the proteins have repetitive structural features.

### Confusing Statistical Significance with Biological Significance

A third common failure pattern is equating statistical significance with biological significance. A high Z-score indicates that the structural similarity is unlikely to have arisen by chance, but it does not indicate that the similarity has biological relevance. Two proteins may share a common structural fold for reasons unrelated to shared ancestry or function, such as convergent evolution to a stable fold or shared structural constraints from their biological environment.

The biological significance of a DALI hit requires additional evidence beyond the Z-score. Researchers should consider the functional annotations of the matched proteins, the conservation of functionally important residues, the taxonomic distribution of the proteins, and any experimental data that support a biological relationship.

The distinction between statistical and biological significance is particularly important in large-scale structural genomics projects where many DALI hits are generated automatically. A high Z-score is a starting point for biological investigation, not a conclusion about biological relationship.

## Quality Controls and Reproducibility in DALI Analysis

### Documenting Analysis Parameters

Reproducible DALI analysis requires complete documentation of the analysis parameters and the software environment. The DALI server has configurable parameters that affect the alignment and the resulting Z-scores, including the structural alphabet settings, the minimum alignment length, and the database selection.

Researchers should record the exact parameters used for each DALI search, including the software version and the database version. The nf-core documentation provides community pipeline standards for usage, configuration, and reproducible workflow context, and similar documentation standards should be applied to structural bioinformatics analyses.

The documentation should be stored with the analysis results so that other researchers can reproduce the analysis or understand its limitations. The Carpentries lessons provide foundational training in computing, data, shell, Git, and programming that supports reproducible research practices, and these skills are directly applicable to structural bioinformatics workflows.

### Validating Results with Independent Methods

DALI Z-scores should be validated with independent structural comparison methods before drawing strong conclusions. The TM score and the GDT_TS score provide complementary measures of structural similarity that can confirm or challenge the DALI results. The random model approach demonstrates that probabilities estimated from the method correlate well with common measures of structural similarity, such as the DALI Z-score and the GDT_TS score, which supports the use of multiple metrics for validation.

Researchers should also consider visual inspection of the structural superposition. The DALI output includes structural alignments that can be visualized in molecular graphics software, and visual inspection can reveal alignment artifacts or biologically meaningful features that are not captured by numerical scores.

For high-stakes conclusions, such as functional annotation transfer or evolutionary inference, researchers should validate the DALI results with additional evidence from sequence analysis, phylogenetic analysis, or experimental studies. The NCBI Data Resources provide access to sequence databases and analysis tools that can support this validation.

### Managing Database Version Effects

DALI Z-scores depend on the database version used for the search. When the database is updated with new structures, the background distribution of random comparisons changes, and the Z-scores for a given query can shift. Researchers should record the database version and repeat the search with the same version when comparing results across different queries or different time points.

For longitudinal studies that track structural similarity over time, researchers should either use a fixed database version or account for database version effects in their analysis. The Galaxy Training Network provides accessible workflow training and analysis tutorials that emphasize reproducibility context, and these principles apply to managing database version effects in structural bioinformatics.

The practical consequence of database version effects is that Z-scores from different searches are not directly comparable unless the database version is the same. Researchers should archive the database version with their results and use consistent versions within a research project.

## Limitations of DALI Z-Scores

### Structural Coverage and Alignment Artifacts

DALI Z-scores are limited by the structural coverage of the alignment. Proteins with low structural coverage, where only a small fraction of residues are aligned, may produce misleading Z-scores. The Z-score reflects the statistical significance of the aligned regions, but it does not indicate what fraction of the protein is involved in the similarity.

Alignment artifacts can also affect Z-scores. Proteins with repetitive structural features, such as beta-propellers or alpha-solenoids, can produce alignments that are geometrically valid but biologically misleading. The DALI algorithm may align different repeats of the same protein in ways that inflate the Z-score without indicating genuine evolutionary relatedness.

Researchers should examine the alignment details, including the aligned residue ranges and the structural superposition, to identify potential artifacts. The root mean square deviation and the alignment length provide context for interpreting the Z-score, and unusual patterns should trigger closer inspection.

### Size and Shape Dependence

The random model approach emphasizes that significance assessment must account for molecular size and shape. The random model restricts comparisons to molecules of the same size and shape as the superposition of interest, and this restriction affects the expected distribution of scores.

The size and shape dependence means that Z-scores from comparisons of proteins with very different sizes or shapes may not be directly comparable. A Z-score of 5.0 for a comparison of two 100-residue domains may have different practical implications than a Z-score of 5.0 for a comparison of two 500-residue domains.

Researchers should consider the size and shape of their query protein when interpreting Z-scores. If the query protein is unusually large or small relative to the database population, the standard threshold categories may need adjustment. The p-value calculation approach, which depends on root mean square deviation, radius of gyration, and thinnest molecular dimension, provides a size-aware alternative to raw Z-score thresholds.

### Multiple Testing and Database Search Context

DALI database searches involve thousands of pairwise comparisons, and the multiple testing problem affects the interpretation of Z-scores. A Z-score that is significant in a single pairwise comparison may not be significant when considered among thousands of comparisons.

Researchers should apply multiple testing corrections or use empirical thresholds based on the distribution of Z-scores across all hits in a search. The false discovery rate approach, which controls the expected proportion of false positives among the hits declared significant, is appropriate for DALI database searches.

The practical consequence is that the threshold for confident significance should be higher in database searches than in single pairwise comparisons. A Z-score of 4.0 may be appropriate for a single pairwise comparison, but a Z-score of 6.0 or higher may be needed for confident significance in a large database search.

## Professional Escalation Criteria for DALI Z-Score Interpretation

### When to Seek Expert Structural Biology Consultation

Researchers should seek expert consultation when DALI results are ambiguous or when the interpretation has high-stakes consequences. Ambiguous results include Z-scores in the 2.0 to 4.0 range with conflicting evidence from other metrics, alignments with unusual patterns, or results that contradict established structural classifications.

Expert consultation is also appropriate when the DALI results will be used for regulatory submissions, clinical decisions, or high-impact publications. Structural biology experts can provide guidance on the appropriate statistical methods, the interpretation of alignment details, and the limitations of the analysis.

The EMBL-EBI Training provides bioinformatics learning pathways, data-resource training, and practical analysis education that can help researchers build the skills needed for independent DALI interpretation. Researchers who encounter repeated difficulties with DALI interpretation should consider formal training in structural bioinformatics.

### When to Repeat or Redesign the Analysis

Researchers should repeat the DALI analysis when the input structures are updated, when the database version changes, or when the analysis parameters are modified. The Z-scores from the original analysis may not be valid after these changes, and the analysis should be repeated to ensure that the conclusions are based on current data.

Researchers should redesign the analysis when the query protein has unusual properties that are not well handled by the standard DALI workflow. Unusual properties include very large or very small proteins, proteins with extensive disorder, proteins with repetitive structural features, or proteins with multiple domains that may align independently.

The redesign may involve using different structural comparison methods, adjusting the alignment parameters, or using a curated database of structurally characterized proteins. The Galaxy Training Network provides accessible workflow training and analysis tutorials that can support the design of reproducible structural bioinformatics workflows.

### When to Report Results with Caveats

Researchers should report DALI results with caveats when the Z-scores are marginal, when the alignment quality is poor, or when the biological interpretation is uncertain. The caveats should be specific and actionable, indicating what additional evidence is needed to strengthen the conclusions.

The reporting should include the Z-score, the root mean square deviation, the alignment length, the database version, and the analysis parameters. This information enables other researchers to evaluate the strength of the evidence and to reproduce the analysis if needed.

The NCBI Data Resources provide official descriptions of databases, search systems, sequence resources, and analysis services that can support the reporting and validation of structural bioinformatics results. Researchers should use these resources to provide context for their DALI results and to enable other researchers to access the underlying data.

## Safety and Regulatory Context for Structural Bioinformatics

### Data Integrity and Reproducibility Standards

Structural bioinformatics analyses, including DALI searches, should follow data integrity and reproducibility standards that are consistent with good scientific practice. The nf-core documentation provides community pipeline standards for usage, configuration, and reproducible workflow context, and these standards can be applied to structural bioinformatics workflows.

The reproducibility standards include version control for analysis scripts, documentation of software environments, archiving of input data and results, and validation of results with independent methods. The Carpentries lessons provide foundational training in computing, data, shell, Git, and programming that supports these reproducibility practices.

For research that will be published or used in regulatory submissions, the analysis should be documented in sufficient detail that an independent researcher can reproduce the results. The documentation should include the exact software versions, database versions, analysis parameters, and data processing steps.

### Ethical Use of Structural Data

Structural bioinformatics research should follow ethical standards for data use and sharing. Researchers should respect the terms of use for structural databases, acknowledge the original structural determination efforts, and share their own results in ways that enable further research.

The NCBI Data Resources provide official descriptions of databases, search systems, sequence resources, and analysis services, and researchers should follow the usage policies for these resources. The EMBL-EBI Training provides bioinformatics learning pathways and data-resource training that includes guidance on responsible data use.

Researchers should also consider the potential dual-use implications of structural bioinformatics research. Structural information about pathogens or toxins could have applications in both beneficial and harmful contexts, and researchers should be aware of the relevant regulations and institutional policies.

### Professional Judgment in Interpretation

The interpretation of DALI Z-scores requires professional judgment that goes beyond mechanical threshold application. Researchers should consider the biological context, the quality of the input structures, the alignment details, and the consistency of the results with other evidence.

Professional judgment is particularly important when the DALI results are used to support conclusions about protein function, evolutionary relationships, or clinical relevance. The Z-score provides statistical evidence for structural similarity, but the biological interpretation requires integration of multiple lines of evidence.

Researchers who are uncertain about the interpretation of DALI results should consult with structural biology experts, review the relevant literature, and consider additional validation experiments. The Galaxy Training Network provides accessible workflow training and analysis tutorials that can support the development of interpretation skills.

## Frequently Asked Questions

### What does a DALI Z-score of 2.0 actually mean?

A DALI Z-score of 2.0 means that the observed structural similarity score is two standard deviations above the mean score expected from random protein pairs of comparable size and shape. In a standard normal distribution, approximately 2.3 percent of random comparisons would produce a score this high or higher. This Z-score indicates weak evidence for structural similarity that warrants further investigation but is not sufficient for confident conclusions about evolutionary or functional relationships.

### How is a DALI Z-score different from a raw alignment score?

A raw alignment score reflects the geometric quality of the structural superposition, while a Z-score normalizes that score against the background distribution of random comparisons. The Z-score accounts for the size and shape of the proteins being compared and expresses the significance of the observed similarity in standardized units. Two alignments with the same raw score can have very different Z-scores if the background distributions differ.

### What Z-score should I use as a threshold for confident structural similarity?

The appropriate threshold depends on the research context. For single pairwise comparisons, a Z-score above 4.0 generally indicates confident structural similarity. For database searches involving thousands of comparisons, a higher threshold of 6.0 or above may be needed to account for multiple testing. The threshold should also be adjusted for the size and shape of the query protein relative to the database population.

### Can a high DALI Z-score occur by chance?

Yes, high DALI Z-scores can occur by chance, particularly in large database searches where thousands of comparisons are performed. The Z-score expresses the significance relative to a random model, but the multiple testing problem means that some high Z-scores will arise purely by chance. Researchers should examine the distribution of Z-scores across all hits and apply multiple testing corrections when appropriate.

### How does protein size affect DALI Z-score interpretation?

Protein size affects the background distribution of random structural comparisons. Larger proteins have more opportunities to achieve a given level of structural similarity by chance, so the same Z-score may have different practical implications for proteins of different sizes. The random model approach accounts for size and shape by restricting comparisons to molecules of the same size and shape as the superposition of interest.

### What other metrics should I examine alongside the DALI Z-score?

Researchers should examine the root mean square deviation, the alignment length, the percentage sequence identity, and the GDT_TS or TM score alongside the DALI Z-score. These metrics provide complementary information about the geometric quality of the alignment and the extent of the structural similarity. The random model approach demonstrates that probabilities estimated from the method correlate well with common measures of structural similarity, including the DALI Z-score and the GDT_TS score.

### How should I report DALI Z-scores in publications?

Publications should report the DALI Z-score, the root mean square deviation, the alignment length, the database version, and the software version. The reporting should include the context for interpretation, such as the distribution of Z-scores across all hits in a database search and the thresholds used for significance. This information enables other researchers to evaluate the strength of the evidence and to reproduce the analysis.

### When should I seek expert help with DALI Z-score interpretation?

Seek expert help when the DALI results are ambiguous, when the interpretation has high-stakes consequences, or when the results will be used for regulatory submissions or clinical decisions. Expert help is also appropriate when the query protein has unusual properties, such as very large size, extensive disorder, or repetitive structural features, that may not be well handled by standard DALI workflows.

## Related Bioinformatics Guides

- [How to Interpret Gene Set Enrichment Analysis Results](/knowledge/bioinformatics/how-to-interpret-gene-set-enrichment-analysis-results)
- [Volcano Plot Proteomics: How to Create and Interpret Them Effectively](/knowledge/bioinformatics/volcano-plot-proteomics-how-to-create-and-interpret-them-effectively)
- [Lipidomic Analysis: A Beginner's Guide to Workflows and Data Interpretation](/knowledge/bioinformatics/lipidomic-analysis-a-beginner-s-guide-to-workflows-and-data-interpretation)
- [Pathway Enrichment Analysis for Proteomics: Tools and Interpretation](/knowledge/bioinformatics/pathway-enrichment-analysis-for-proteomics-tools-and-interpretation)
- [Proteomics Analysis Tools: A Comparative Guide for Functional Interpretation](/knowledge/bioinformatics/proteomics-analysis-tools-a-comparative-guide-for-functional-interpretation)

## References and Further Reading

- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Bioconductor](https://bioconductor.org/). Bioconductor Project.
- [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.
- [nf-core Documentation](https://nf-co.re/docs). nf-core.
- [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.
- [Differentiating tumor recurrence and pseudoprogression in postoperative gliomas using pseudo-continuous arterial spin labeling (pCASL) technique.](https://pubmed.ncbi.nlm.nih.gov/41419837). BMC medical imaging, 2025.
- [Statistics of random protein superpositions: p-values for pairwise structure alignment.](https://pubmed.ncbi.nlm.nih.gov/18333756). Journal of computational biology : a journal of computational molecular cell biology, 2008.
- [Effects of vitamin A and vitamin D&lt,sub&gt,3&lt,/sub&gt, supplementation on child growth and development in low- and middle-income countries: a systematic review and meta-analysis.](https://doi.org/10.21037/tp-2025-507). 2025.
- [Machine learning-based identification of key genes underlying sex differences in hepatocellular carcinoma and targeted drug screening.](https://doi.org/10.3892/br.2026.2147). 2026.
- [Analysis of the Epidemiological Characteristics and Risk Factors for Severe Progression of Scrub Typhus in the Dali Region of China.](https://doi.org/10.12659/msm.951111). 2026.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.