# Common Mistakes in Domain Annotation: Why Your Protein Domain Boundaries Are Wrong and How to Fix Them

Protein domain annotation assigns discrete structural and evolutionary units to specific regions of an amino acid sequence. When boundaries are placed incorrectly, downstream functional conclusions about active sites, binding interfaces, and disease variants become unreliable. This article explains why sequence-based domain annotation frequently produces incorrect boundaries, how structural information from deep learning models can correct those errors, and what practical steps you can take to validate and refine domain assignments in your own research.

The core problem is straightforward: most domain annotation tools rely on hidden Markov models (HMMs) trained on sequence motifs, and these models lack access to the three-dimensional geometry that actually defines a domain. Sequence-based methods are prone to error, particularly in calling domain boundaries and the motifs within them, because they are limited by a lack of structural information accessible to the model. With the advent of deep learning-based protein structure prediction, existing sequence-based domain annotation methods can be improved by taking into account the geometry of protein structures. This article provides a practical framework for recognizing when your domain boundaries are wrong, using structural evidence to correct them, and documenting your decisions so that functional interpretations rest on solid ground.

## At a Glance: Domain Annotation Decision Framework

The table below summarizes the most common annotation scenarios, the typical failure mode, and the corrective action you should take before proceeding with functional analysis.

| Annotation Scenario | Common Failure Mode | Recommended Corrective Action |
| --- | --- | --- |
| Single-domain protein annotated by sequence HMM alone | Boundary overextension into disordered termini or linker regions | Compare against predicted structure and check inter-residue distance patterns |
| Multi-domain protein with tandem repeats | Repeat units merged or split incorrectly due to motif degeneracy | Use structure-based repeat detection and dimensionality reduction on geometric features |
| Novel fold with no sequence homologs | Domain boundaries assigned by homology transfer from distant relatives | Use structure-based domain parsing tools that integrate predicted aligned errors |
| Truncated sequence or partial model | Domain appears incomplete or boundaries fall inside a packed core | Check for sequence coverage issues and re-run annotation on the full-length construct |
| Tightly packed domain duplications | Two domains misidentified as one due to close packing | Examine predicted aligned error maps for internal domain boundaries |

## Why Sequence-Based Annotation Produces Wrong Boundaries

### The Hidden Markov Model Limitation

Most domain annotation pipelines begin with profile-based searches against databases such as Pfam, CDD, or InterPro. These tools use HMMs that capture conserved sequence patterns across a family of related proteins. The strength of this approach is speed and scalability: you can annotate an entire proteome in minutes. The weakness is that HMMs are trained on sequence motifs, and they have no direct access to the three-dimensional structure that physically defines a domain.

The practical consequence is that HMM-based annotations frequently place boundaries at positions that do not correspond to the actual structural unit. A domain boundary in structural terms is a region where the polypeptide chain folds back on itself to create a compact globular unit, often with a hydrophobic core. Sequence-based methods infer these boundaries from conservation patterns and insertion-deletion histories, which are imperfect proxies for physical compactness. When a protein contains long disordered regions, flexible linkers, or inserted domains, sequence-based tools often extend or truncate boundaries in ways that do not match the structural reality.

### The Structural Information Gap

The fundamental limitation of sequence-based annotation is the lack of structural information accessible to the model. A domain is a physical object with specific geometry, packing interactions, and surface properties. Sequence motifs capture evolutionary conservation, but they do not capture the spatial relationships that define a domain boundary. For example, a domain may contain a conserved motif that is essential for function, but the actual boundary of the domain may extend well beyond the motif into regions that are structurally integral but poorly conserved.

This gap becomes especially problematic for repeat proteins. Leucine-rich repeat (LRR) domains, for example, consist of tandem arrays of short structural units that form a solenoid shape. Each repeat unit is typically 20 to 30 residues, and the boundaries between units are defined by the geometry of the beta-alpha structural motif. Sequence-based HMMs often struggle with these boundaries because the individual repeats are degenerate: they share a common structural fold but have highly variable sequences. The result is that repeat units are frequently merged, split, or shifted relative to the true structural units.

### The Scale of the Problem

The release of hundreds of millions of high-accuracy predicted structures has changed the scale of the annotation problem. The AlphaFold Database contains approximately 200 million models, and partitioning these models into domains and assigning them to an evolutionary hierarchy is an efficient way to gain functional insights into proteins. However, classifying such a large number of predicted structures challenges the infrastructure of current structure classifications. Manual curation is impossible at this scale, and automated tools must be able to recognize domains from structural geometry instead of relying solely on sequence conservation.

The challenge is also academic. Incorrect domain boundaries lead to incorrect functional conclusions. If you assign a functional motif to the wrong domain, or if you truncate a domain such that a binding interface is excluded, your interpretation of experimental results will be wrong. This is particularly dangerous in clinical contexts, where domain boundaries are used to interpret the pathogenicity of genetic variants.

## Core Principles of Structure-Aware Domain Annotation

### Domains Are Physical Objects

The first principle of correct domain annotation is that domains are physical objects defined by their three-dimensional structure. A domain is a compact, independently folding unit of a protein, typically with a hydrophobic core and a defined surface. The boundaries of a domain are the points where the polypeptide chain exits the compact unit and enters a linker, a disordered region, or another domain.

This physical definition has practical consequences. The boundary of a domain is not necessarily at the edge of a conserved sequence motif. It is at the point where the chain transitions from compact packing to extended conformation. Sequence-based tools cannot see this transition directly, but structure-based tools can.

### Predicted Structures Provide Geometric Evidence

Deep learning-based structure prediction methods such as AlphaFold produce models with near-atomic accuracy for a large fraction of proteins. These models provide geometric evidence that can be used to correct sequence-based annotations. The key insight is that the geometry of a protein structure contains information about domain boundaries that is not present in the sequence alone.

Structure-based domain parsing works by examining inter-residue distances in three-dimensional structures. A domain boundary is typically a region where the chain makes a sharp turn or where the packing density drops. By analyzing the matrix of inter-residue distances, automated tools can identify the compact globular units that correspond to domains.

### Predicted Aligned Errors Reveal Domain Organization

One of the most useful outputs of AlphaFold-style prediction is the predicted aligned error (PAE) matrix. This matrix estimates the expected positional error between pairs of residues when the model is aligned on one of them. The PAE matrix has a characteristic block-like structure for multi-domain proteins: residues within the same domain have low predicted aligned error relative to each other, while residues in different domains have high predicted aligned error.

This PAE structure is a powerful signal for domain boundary detection. Automated domain parsers use PAE information to identify the boundaries between domains, even when the domains are tightly packed or when sequence conservation is weak. If you are annotating a protein manually, examining the PAE plot should be one of your first steps.

## Practical Workflow for Correcting Domain Boundaries

### Step 1: Generate or Retrieve a Predicted Structure

Before you finalize any domain annotation, you need a structural model. If your protein is in the AlphaFold Database, retrieve the model directly. If not, run a structure prediction using an appropriate tool. The quality of the model matters: check the predicted local distance difference test (pLDDT) scores to identify regions of low confidence, and be cautious about annotating domains in regions where the model is unreliable.

### Step 2: Run Sequence-Based Annotation as a Starting Point

Run your standard sequence-based annotation tools first. This gives you a baseline and identifies conserved motifs that are likely to be functionally important. Record the boundaries suggested by each tool, and note where different tools disagree. Disagreement between tools is often the first sign that boundaries are ambiguous.

### Step 3: Examine the Predicted Aligned Error Matrix

Plot the PAE matrix for your protein. Look for the block-like structure that indicates domain organization. Low PAE blocks correspond to residues that are structurally coupled, which typically means they are in the same domain. High PAE regions between blocks indicate domain boundaries. Mark the approximate boundary positions suggested by the PAE structure.

### Step 4: Analyze Inter-Residue Distances and Packing

Examine the three-dimensional structure directly. Look for the compact globular units that define domains. Check whether the sequence-based boundaries fall at the edges of these compact units or whether they cut through the middle of a packed core. A boundary that falls inside a packed core is almost certainly wrong.

For repeat proteins, use dimensionality reduction methods on the geometric features of the structure to identify individual repeat units. These methods can correct mistakes made by existing machine learning-based annotation tools and enable the automated detection of hairpin loops and structural anomalies in the solenoid.

### Step 5: Compare Against Known Domain Classifications

If your protein has homologs in domain classification databases such as ECOD, compare your boundaries against the evolutionary classification. Homology-based domain assignment using sequence and structural similarity searches can identify domains that are missed by de novo parsing. However, be aware that homology-based methods inherit the errors of their training data, and they may miss novel domains that have no known relatives.

### Step 6: Validate with Multiple Lines of Evidence

A confident domain annotation should be supported by multiple independent lines of evidence. Sequence conservation, structural compactness, PAE patterns, and evolutionary classification should all point to the same boundaries. If these lines of evidence disagree, investigate the source of the disagreement before proceeding.

### Step 7: Document Your Decisions

Record the final domain boundaries, the evidence supporting them, and any disagreements with sequence-based tools. This documentation is essential for reproducibility and for interpreting downstream functional experiments. If you publish your results, include the domain boundaries and the method used to determine them.

## Tools and Options for Structure-Based Domain Parsing

### Automated Domain Parsers

Several automated tools have been developed to parse domains from predicted structures. These tools use different combinations of structural and evolutionary information to identify domain boundaries.

One approach is represented by DPAM, a Domain Parser for AlphaFold Models. DPAM automatically recognizes globular domains from predicted structures based on inter-residue distances in three-dimensional structures, predicted aligned errors, and ECOD domains found by sequence and structural similarity searches. In a benchmark of 18,759 AlphaFold models, DPAM recognized 98.8% of domains and assigned correct boundaries for 87.5%, significantly outperforming structure-based domain parsers and homology-based domain assignment using ECOD domains found by sequence or structural searches.

The practical implication is that automated structure-based parsing is now reliable enough to use as a primary annotation tool, at least for well-folded globular domains. However, the 87.5% boundary accuracy means that roughly one in eight domains will have incorrect boundaries even with the best automated tools. Manual inspection remains necessary for critical applications.

### Dimensionality Reduction for Repeat Proteins

For repeat proteins, standard domain parsers often fail because the individual repeat units are not independently folding domains. They are structural units within a larger solenoid domain, and their boundaries are defined by the geometry of the repeat array instead of by packing transitions.

Dimensionality reduction methods can annotate repeat units by analyzing the geometry of the protein structure. These methods map the structural features of each residue into a low-dimensional space where the repeat units appear as distinct clusters. The boundaries between clusters correspond to the boundaries between repeat units. These methods can correct mistakes made by existing machine learning-based annotation tools and enable the automated detection of hairpin loops and structural anomalies in the solenoid.

### Homology-Based Assignment

Homology-based domain assignment uses sequence and structural similarity searches to identify domains that match known evolutionary families. This approach is powerful when your protein has clear homologs, but it fails for novel domains that lack sequence or structural similarity to known folds.

The limitations of homology-based assignment are particularly evident for candidate novel folds. Many candidate novel fold domains lack clear sequence or structural similarity to known folds, and some appear as insertions into known transmembrane or enzymatic domains. Others occur in modular architectures, co-occurring with interaction or catalytic folds such as beta-barrels, zinc fingers, or Rossmann-like domains. Homology-based tools will miss these domains entirely, and de novo structural parsing is required.

### Choosing Between Tools

The choice of tool depends on your specific question. If you need to annotate a large number of proteins and are willing to accept some boundary errors, an automated parser such as DPAM is appropriate. If you are working on a single protein of high importance, manual inspection of the structure and PAE matrix is warranted. If you are working with repeat proteins, dimensionality reduction methods are likely to outperform generic domain parsers.

The best practice is to use multiple tools and compare their outputs. Where they agree, you can be confident in the annotation. Where they disagree, investigate the structural evidence to determine which annotation is correct.

## Records and Measurements for Domain Annotation

### What to Record

For each protein you annotate, record the following information:

- The sequence identifier and version
- The source of the structural model (database ID or prediction run)
- The model quality metrics (pLDDT scores, particularly in boundary regions)
- The boundaries suggested by each annotation tool
- The final boundaries and the evidence supporting them
- Any disagreements between tools and how they were resolved
- The date and version of all software used

This record is essential for reproducibility. If you later discover that a boundary was wrong, you need to be able to trace which tool produced the error and why.

### Quality Metrics for Boundary Confidence

Several metrics can help you assess the confidence of a domain boundary:

- The pLDDT score at the boundary region. Low pLDDT indicates low model confidence, and boundaries in low-confidence regions should be treated with caution.
- The PAE values between residues on either side of the proposed boundary. High PAE across the boundary supports the existence of a domain boundary.
- The packing density of the residues near the boundary. A boundary that cuts through a densely packed core is suspect.
- The agreement between independent annotation methods. Agreement increases confidence.

### Benchmarking Against Manually Annotated Data

If you are developing or evaluating an annotation method, benchmark against manually annotated domain data. For example, structure-aware annotation methods for LRR domains have been validated against a benchmark dataset of 172 manually annotated LRR domains. Manual annotation by experts remains the gold standard, and automated methods should be evaluated against it before being used for high-stakes applications.

## Common Failure Patterns in Domain Annotation

### Boundary Overextension into Linkers

The most common failure pattern is extending a domain boundary into a flexible linker or disordered region. Sequence-based tools often include conserved residues that lie outside the compact structural unit, particularly when the linker contains short conserved motifs. The result is a domain annotation that includes residues that are not part of the structural domain.

The fix is to examine the structure and identify the point where the chain transitions from compact packing to extended conformation. The domain boundary should be placed at this transition point, not at the edge of the conserved sequence motif.

### Boundary Truncation Inside the Domain Core

The opposite failure is truncating a domain such that the boundary falls inside the compact core. This happens when sequence-based tools identify a conserved motif and place the boundary at the motif edge, even though the structural domain extends beyond the motif. The result is a domain annotation that excludes structurally integral residues.

This failure is particularly dangerous because it can exclude residues that are essential for function. If you are interpreting the effects of variants, a truncated domain annotation can cause you to miss pathogenic variants that fall in the excluded region.

### Merging of Tandem Repeats

For repeat proteins, the most common failure is merging multiple repeat units into a single domain. Sequence-based HMMs often fail to resolve the boundaries between individual repeats because the repeats are degenerate in sequence. The result is an annotation that treats the entire solenoid as a single unit, obscuring the modular structure.

Structure-based methods that analyze the geometry of the solenoid can resolve individual repeat units and correct these errors. Dimensionality reduction methods are particularly effective for this purpose.

### Splitting of a Single Domain

The opposite failure is splitting a single domain into multiple domains. This happens when a domain contains a deep cleft or a pronounced surface feature that causes structure-based parsers to identify a false boundary. It can also happen when a domain contains a large insertion that creates a sub-structure with distinct packing.

The fix is to examine the PAE matrix and the evolutionary classification. If the region in question has low PAE throughout and belongs to a single evolutionary family, it is likely a single domain despite the surface features.

### Misassignment of Inserted Domains

Many proteins contain inserted domains that interrupt the sequence of a larger domain. These insertions are often missed by sequence-based tools, which may annotate the flanking regions as a single domain and ignore the insertion. Structure-based tools can identify insertions because they appear as distinct compact units in the structure.

The practical consequence of missing an inserted domain is that you may misinterpret the architecture of the protein. Inserted domains often have distinct functions, such as binding or catalytic activity, and missing them can lead to incorrect functional conclusions.

### Errors from Truncated Sequences

Domain annotation is only as good as the sequence you provide. If your sequence is truncated, either because of a partial gene model or because you are working with a fragment, the domain boundaries will be wrong. Truncated sequences can also cause errors in structure prediction, leading to incorrect models and incorrect domain annotations.

The fix is to work with full-length sequences whenever possible. If you must work with a fragment, be aware that domain boundaries near the truncation point are unreliable.

## Limitations of Current Methods

### The Boundary Accuracy Ceiling

Even the best automated domain parsers do not achieve perfect boundary accuracy. DPAM, for example, assigns correct boundaries for 87.5% of domains in a large benchmark. This means that for a typical protein with several domains, there is a meaningful probability that at least one boundary is wrong.

The implication is that automated annotation should be treated as a starting point, not a final answer. For high-stakes applications, manual inspection of the structural evidence is essential.

### The Challenge of Novel Folds

Domain annotation is particularly challenging for novel folds that lack sequence or structural similarity to known domains. Many candidate novel fold domains appear as insertions into known transmembrane or enzymatic domains, while others occur in modular architectures with interaction or catalytic folds. Some resemble known folds that have undergone topological rearrangements or circular permutations, while others result from errors in domain boundary prediction, often due to truncated sequences or tightly packed domain duplications.

For these challenging cases, integrating structural, evolutionary, and contextual information is essential. No single method is sufficient, and manual curation may be required.

### The Problem of Low-Confidence Predictions

Structure prediction models are not equally accurate for all proteins. Some regions of a protein may have low pLDDT scores, indicating low model confidence. Domain boundaries in these regions are unreliable, and functional conclusions based on them should be treated with caution.

If your protein has large regions of low confidence, consider whether the predicted structure is suitable for domain annotation at all. In some cases, experimental structure determination may be necessary.

### The Evolutionary Context Gap

Domain annotation is about identifying compact structural units and assigning domains to an evolutionary hierarchy. This evolutionary context provides functional insights that are not available from structure alone. However, evolutionary classification is limited by the coverage of existing databases, and novel domains may not have clear evolutionary relationships.

The integration of structural and evolutionary information is an active area of research. Tools that combine structure-based parsing with homology-based assignment are more accurate than either approach alone, but they still struggle with truly novel domains.

## Safety and Regulatory Context for Clinical Applications

### Domain Boundaries in Variant Interpretation

Domain boundaries play a critical role in the interpretation of genetic variants. When a variant falls within a domain, it is often assumed to affect the function of that domain. When a variant falls outside annotated domains, it may be dismissed as benign. If the domain boundaries are wrong, this logic fails.

A concrete example comes from studies of RELA variants. Monoallelic RELA variants resulting in haploinsufficiency have been linked to recurrent mucocutaneous ulcers and enteritis, while heterozygous RELA dominant-negative variants often exhibit autoinflammatory phenotypes associated with type I interferonopathy. Functional studies using RELA nonsense variants identified amino acid P290 as the boundary between haploinsufficiency and dominant-negative variants in non-missense deleterious variants. The positions of non-missense deleterious variants in RELA allow the estimation of the associated functional changes, facilitating the precise diagnosis of autosomal dominant RelA deficiency.

This example illustrates the clinical importance of precise boundary determination. If the functional boundary between variant classes is at a specific residue, then accurate annotation of that region is essential for clinical decision-making. Conversely, missense variants require functional verification, as their impact cannot be predicted from position alone.

### Reproducibility Requirements

Regulatory submissions and clinical reports increasingly require reproducible bioinformatics analysis. This means that your domain annotation must be documented in sufficient detail that another researcher can reproduce your results. The documentation should include the software versions, the databases used, the parameters set, and the evidence supporting each boundary decision.

Reproducible analysis workflows are supported by platforms such as Galaxy, which provides accessible workflow training and analysis tutorials, and nf-core, which provides community pipeline standards and usage documentation. These platforms can help you build reproducible domain annotation pipelines.

### Professional Escalation Criteria

You should escalate to a structural biology expert or a clinical genetics professional when:

- The domain annotation affects a clinical interpretation, such as variant pathogenicity classification
- Multiple annotation methods disagree on boundaries in a functionally important region
- The predicted structure has large regions of low confidence in a domain boundary region
- The protein contains novel folds or insertions that cannot be confidently annotated
- The functional conclusions from your experiments depend critically on a specific domain boundary

In these situations, the cost of a wrong boundary is high enough to justify expert review.

## Professional Escalation Criteria

### When to Seek Expert Review

Domain annotation errors are not always obvious, and the consequences of errors vary with the application. You should seek expert review when:

- You are interpreting the clinical significance of genetic variants and the variant falls near a domain boundary
- You are designing mutagenesis experiments and the domain boundary determines which residues to mutate
- You are interpreting protein-protein interaction data and the interaction interface falls near a domain boundary
- You are annotating a protein with no close homologs and the domain architecture is uncertain
- You are working with repeat proteins and the repeat boundaries are ambiguous

### What to Prepare for Expert Review

When you escalate to an expert, provide:

- The sequence and the predicted structure
- The PAE matrix and pLDDT scores
- The annotations from all tools you have tried
- The specific question you need answered
- The downstream application that depends on the annotation

This preparation allows the expert to focus on the specific uncertainty instead of redoing your entire analysis.

## Building a Domain Boundary Evidence Log: A Practical Decision Framework

### Why Your Annotation Needs a Structured Record

Most domain annotation errors are not caught at the time they are made. They surface months later when a functional experiment produces unexpected results or when a collaborator questions a boundary in a manuscript revision. Without a structured record of how you arrived at your boundaries, you cannot trace the source of the error or determine whether the annotation or the experiment is wrong. A domain boundary evidence log solves this problem by forcing you to document the reasoning behind each boundary decision at the time you make it.

The evidence log serves three distinct purposes. First, it creates a reproducible trail that another researcher can follow to understand why you placed boundaries where you did. Second, it provides a mechanism for detecting annotation errors early by making disagreements between evidence types visible. Third, it gives you a defensible basis for boundary decisions when reviewers or collaborators question them. The log is not busywork. It is the practical mechanism that converts domain annotation from an opaque process into a transparent one.

### The Five-Column Boundary Decision Table

For each domain boundary in your protein, create a row in a structured table with five columns: boundary position, sequence evidence, structural evidence, evolutionary evidence, and confidence rating. This table becomes the core of your evidence log and forces you to articulate what supports each boundary.

The boundary position column records the residue numbers you have assigned as the start and end of the domain. The sequence evidence column documents what your HMM-based tools reported, including the specific tool names and the boundaries they suggested. The structural evidence column records what the predicted structure shows, including the pLDDT scores at the boundary region, the PAE values across the proposed boundary, and whether the boundary falls at the edge of a compact globular unit or cuts through a packed core. The evolutionary evidence column documents whether the boundary matches known domain classifications from databases such as ECOD or Pfam, and whether homologous proteins show similar boundary placements. The confidence rating column assigns a qualitative score of high, medium, or low based on how well the evidence types agree.

The act of filling in this table reveals problems immediately. If your sequence evidence places a boundary at residue 150 but your structural evidence shows that residue 150 sits in the middle of a densely packed hydrophobic core, you have a conflict that needs resolution before you proceed. If your structural evidence supports a boundary but your evolutionary evidence shows that all known homologs place the boundary elsewhere, you need to investigate whether you are looking at a genuine novel feature or an annotation artifact.

### Assigning Confidence Ratings with Clear Criteria

Confidence ratings must be tied to explicit criteria so that different researchers applying the framework to the same protein reach the same conclusion. A high confidence rating requires that all three evidence types agree on the boundary position within a tolerance of five residues, that the pLDDT score at the boundary region is above the reliability threshold for your prediction tool, and that the PAE values across the boundary are clearly elevated relative to values within the domain.

A medium confidence rating applies when two evidence types agree but the third is ambiguous or unavailable. For example, you may have strong structural evidence for a boundary but weak sequence evidence because your protein has no close homologs in the HMM databases. A medium rating is also appropriate when the evidence types agree on the general region but disagree on the exact position by more than five residues.

A low confidence rating applies when evidence types conflict, when the boundary falls in a region of low model confidence, or when you are annotating a novel fold with no evolutionary context. Low confidence boundaries require special handling. You should not use them as the basis for functional conclusions without additional validation, and you should flag them for expert review if they affect a high-stakes decision.

### The Disagreement Resolution Protocol

When evidence types disagree, follow a structured protocol instead of relying on intuition. The protocol has four steps. First, document the disagreement in your evidence log, recording exactly what each evidence type suggests and where the conflict lies. Second, examine the structural evidence more closely, because structure is the most direct physical evidence of domain organization. Look at the packing density around the disputed boundary, the PAE values, and whether the proposed boundary cuts through secondary structure elements. Third, check whether the disagreement stems from a known failure mode. For example, sequence tools often overextend boundaries into linkers, and structure-based parsers sometimes split a single domain at a deep surface cleft. Fourth, if the disagreement persists after this analysis, assign a medium or low confidence rating and document the uncertainty instead of forcing a resolution.

The protocol prevents two common errors. The first error is defaulting to the sequence-based annotation because it is familiar and easy to defend. The second error is defaulting to the structure-based annotation because it feels more modern. Neither default is correct. The correct response is to investigate the source of the disagreement and to document what you find.

### Recording Model Quality Metrics at Boundary Regions

The reliability of your structural evidence depends on the quality of the predicted model at the boundary region. A boundary placed in a region of high pLDDT is on firmer ground than a boundary placed in a region of low pLDDT. Your evidence log should record the pLDDT score at each boundary and note whether the boundary falls in a confidently predicted region.

For AlphaFold-style models, the pLDDT score provides a per-residue estimate of model confidence. Scores above 90 indicate high confidence, scores between 70 and 90 indicate confident prediction, and scores below 50 indicate low confidence. Boundaries in low-confidence regions should receive a lower confidence rating regardless of how well the other evidence types agree, because the structural evidence itself is unreliable.

The PAE matrix provides complementary information about the relative positional confidence between residue pairs. For domain annotation, the key pattern is the block-like structure that indicates domain organization. Residues within the same domain have low PAE values relative to each other, while residues in different domains have high PAE values. Your evidence log should record whether the PAE pattern supports the proposed boundary and whether the transition between low and high PAE regions is sharp or gradual. A gradual transition suggests that the boundary is not well defined by the structural model.

### The Boundary Revision Trigger List

Your evidence log should also include a list of triggers that prompt you to revisit and potentially revise your domain boundaries. These triggers include new experimental data that conflicts with your functional interpretation, the release of an updated structural model with improved confidence, the identification of a new homolog that changes the evolutionary context, and the discovery of a variant of interest that falls near a boundary.

When a trigger occurs, do not simply revise the boundary and move on. Return to your evidence log, review the original evidence, and determine whether the new information changes the balance of evidence. If it does, update the log with the new boundary, the new evidence, and the reason for the revision. This creates a complete history of how your annotation evolved and why.

### Common Failure Patterns in the Evidence Log Process

The evidence log process fails in predictable ways. The most common failure is filling in the log after the annotation is complete instead of during the annotation process. A retrospective log is little better than no log at all, because it records what you decided but not why you decided it. The evidence that informed your decision is often lost by the time you sit down to document it.

The second failure pattern is treating the confidence rating as a fixed property of the boundary instead of a reflection of the current evidence. A boundary that receives a high confidence rating today may deserve a lower rating tomorrow if new evidence emerges that conflicts with the original annotation. The confidence rating should be updated whenever the evidence changes.

The third failure pattern is using the evidence log to defend a predetermined conclusion instead of to test it. If you have already decided where the boundary should be and you fill in the log to support that decision, you have defeated the purpose of the exercise. The log is most valuable when it reveals problems with your initial assumptions.

### Implementing the Evidence Log in Your Workflow

The evidence log does not require specialized software. A spreadsheet with the five columns described above is sufficient for most projects. For larger annotation efforts involving many proteins, consider using a structured format that can be queried and shared with collaborators. The key requirement is consistency. Every boundary in every protein you annotate should have a row in the log, and every row should be filled in completely before you use the annotation for functional interpretation.

The time cost of maintaining the evidence log is modest relative to the cost of a wrong boundary. A wrong boundary can lead to wasted experiments, incorrect variant interpretations, and manuscript revisions. The evidence log catches many of these errors before they propagate into downstream analysis.

### Professional Escalation Criteria for Boundary Uncertainty

The evidence log provides a clear basis for deciding when to escalate to an expert. Escalate when a boundary receives a low confidence rating and the annotation affects a clinical interpretation, an experimental design, or a published conclusion. Escalate when evidence types conflict and the disagreement resolution protocol does not produce a satisfactory answer. Escalate when you are annotating a novel fold with no evolutionary context and the boundary placement will be used to guide functional experiments.

When you escalate, provide the expert with your complete evidence log for the disputed boundary. This allows the expert to see exactly what evidence you considered and where the uncertainty lies. An expert who receives a well-documented evidence log can focus on resolving the specific uncertainty instead of redoing your entire analysis from scratch.

### Integrating the Evidence Log with Reproducible Workflows

The evidence log complements reproducible analysis workflows by documenting the interpretive decisions that automated pipelines do not capture. Platforms such as Galaxy provide accessible workflow training and analysis tutorials that help you build reproducible annotation pipelines. The nf-core documentation provides community pipeline standards and usage guidance for reproducible analysis. The Carpentries lessons provide foundational computing and data skills that support good record-keeping practices.

The evidence log fills the gap between the computational pipeline and the scientific interpretation. The pipeline produces the raw annotations. The evidence log records how you evaluated those annotations, resolved disagreements, and arrived at final boundaries. Both are necessary for a fully reproducible analysis.

### The Relationship Between Evidence Logs and Domain Classification Databases

Your evidence log should reference the domain classification databases you consulted and record how your boundaries compare to the classifications in those databases. This comparison is particularly important when your boundaries differ from established classifications. A difference is not necessarily an error, but it requires justification. The evidence log provides the space to record that justification.

For example, if you are annotating a protein with a candidate novel fold that lacks clear sequence or structural similarity to known folds, your evidence log should record that the evolutionary evidence column is empty or marked as not applicable. This explicit acknowledgment of missing evidence is more useful than silently omitting the evolutionary comparison, because it flags the boundary as lower confidence and prompts additional scrutiny.

### Using the Evidence Log for Retrospective Analysis

The evidence log becomes increasingly valuable as your annotation portfolio grows. When you annotate a new protein that shares sequence or structural similarity with a previously annotated protein, you can consult the earlier evidence log to see how you handled similar boundary questions. This consistency check prevents you from making different decisions for similar proteins without a documented reason.

The evidence log also supports retrospective error analysis. If a downstream experiment reveals that a domain boundary was wrong, the evidence log tells you which evidence type led you astray. This information helps you calibrate your confidence in different evidence types and improve your annotation process over time.

## Frequently Asked Questions

### Why do sequence-based domain annotation tools give different boundaries for the same protein?

Sequence-based tools use different HMMs, different training data, and different algorithms for boundary placement. Each tool has its own biases, and these biases lead to different boundary predictions. Disagreement between tools is a signal that the boundaries are ambiguous and that structural evidence should be used to resolve the ambiguity.

### How can I tell if a domain boundary is wrong?

The most reliable way is to examine the predicted structure. A correct boundary should fall at the edge of a compact globular unit, not through the middle of a packed core. The PAE matrix should show high predicted aligned error across the boundary and low error within each domain. If the sequence-based boundary does not match the structural evidence, the boundary is likely wrong.

### What is the predicted aligned error matrix and how do I use it for domain annotation?

The predicted aligned error matrix estimates the expected positional error between pairs of residues when the model is aligned on one of them. For multi-domain proteins, the matrix has a block-like structure: residues within the same domain have low error relative to each other, while residues in different domains have high error. You can use this structure to identify domain boundaries by looking for the transitions between low-error blocks.

### Can I rely on automated domain parsers for my final annotation?

Automated parsers are accurate enough for many applications, but they do not achieve perfect boundary accuracy. In a large benchmark, the best parser assigned correct boundaries for 87.5% of domains. For high-stakes applications, such as clinical variant interpretation or experimental design, you should manually inspect the structural evidence and confirm the automated annotation.

### How do I annotate repeat proteins like leucine-rich repeat domains?

Repeat proteins require specialized methods because the individual repeat units are not independently folding domains. Standard domain parsers often fail to resolve repeat boundaries. Dimensionality reduction methods that analyze the geometry of the protein structure can annotate repeat units and correct mistakes made by sequence-based tools. These methods can also detect hairpin loops and structural anomalies in the solenoid.

### What should I do if my protein has a novel fold with no sequence homologs?

Novel folds require integration of multiple lines of evidence. Structure-based parsing can identify compact globular units, but evolutionary classification may be impossible if there are no known relatives. Examine the PAE matrix, the packing density, and the structural context. If the domain appears as an insertion into a known domain, consider whether it has a distinct function. Be prepared to classify the domain as a domain of unknown function if no functional evidence is available.

### How do domain annotation errors affect the interpretation of genetic variants?

If a variant falls within an annotated domain, it is often assumed to affect the function of that domain. If the domain boundary is wrong, a variant that actually falls inside a domain may be classified as outside a functional region. This can lead to incorrect pathogenicity classifications. The RELA example shows that precise boundary determination can be clinically significant, with different variant classes separated by a single amino acid position.

### What records should I keep for reproducible domain annotation?

Record the sequence identifier, the structural model source and quality metrics, the boundaries from each annotation tool, the final boundaries and the evidence supporting them, and the versions of all software and databases used. This documentation allows another researcher to reproduce your analysis and allows you to trace errors if they are discovered later.

## Related Bioinformatics Guides

- [Functional Annotation of Metagenomes: A Guide to Databases and Pipelines](/knowledge/bioinformatics/functional-annotation-of-metagenomes-a-guide-to-databases-and-pipelines)
- [Hidden Markov Models for Protein Domain Annotation](/knowledge/bioinformatics/hidden-markov-models-for-protein-domain-annotation)
- [Functional Metagenomics: From Gene Prediction to Pathway Reconstruction](/knowledge/bioinformatics/functional-metagenomics-from-gene-prediction-to-pathway-reconstruction)
- [Single-Cell Annotation: A Workflow for Cell Type Identification](/knowledge/bioinformatics/single-cell-annotation-a-workflow-for-cell-type-identification)
- [Genomic Prediction in Livestock: A Decision Framework for Breeders](/knowledge/bioinformatics/genomic-prediction-in-livestock-a-decision-framework-for-breeders)

## References and Further Reading

- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Bioconductor](https://bioconductor.org/). Bioconductor Project.
- [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.
- [nf-core Documentation](https://nf-co.re/docs). nf-core.
- [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.
- [Structure-aware annotation of leucine-rich repeat domains.](https://pubmed.ncbi.nlm.nih.gov/39499733). PLoS computational biology, 2024.
- [DPAM: A domain parser for AlphaFold models.](https://pubmed.ncbi.nlm.nih.gov/36539305). Protein science : a publication of the Protein Society, 2023.
- [Exploring the Terra incognita of AI-based domain classifications.](https://pubmed.ncbi.nlm.nih.gov/41288334). Protein science : a publication of the Protein Society, 2025.
- [Discovering patterns in the pathologic significance of non-missense deleterious variants in RELA.](https://pubmed.ncbi.nlm.nih.gov/41823919). The Journal of allergy and clinical immunology, 2026.
- [Structure-Aware Annotation of Leucine-rich Repeat Domains.](https://pubmed.ncbi.nlm.nih.gov/37961157). bioRxiv : the preprint server for biology, 2023.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.