# SCOP vs. CATH: How to Choose the Right Protein Classification Database for Your Research

Researchers who need to classify protein structures face a practical decision between SCOP and CATH, two hierarchical databases that organize the Protein Data Bank by evolutionary and structural relationships. The two systems differ in how they define domains, how they build their hierarchies, and how often they release updates. This article compares SCOP and CATH across classification logic, curation methods, coverage, and practical use cases so you can select the appropriate database for your specific research question, whether you are benchmarking structure comparison tools, training machine learning classifiers, or interpreting the evolutionary relationships of a newly solved structure.

## At a Glance

The table below summarizes the key operational differences between SCOP and CATH that affect database selection.

| Feature | SCOP | CATH |
| --- | --- | --- |
| Classification approach | Purely manual curation by expert visual inspection | Combination of automated methods and manual curation |
| Primary hierarchy levels | Class, Fold, Superfamily, Family | Class, Architecture, Topology (fold), Homologous superfamily |
| Domain definition emphasis | Evolutionary relationships and structural similarity | Structural similarity with evolutionary relationships at the superfamily level |
| Update frequency | Periodic releases with manual curation | More frequent automated updates supplemented by manual review |
| Best suited for | Deep evolutionary analysis, detailed manual curation, structural phylogenetics | Large-scale automated classification, genome annotation, integration with InterPro |
| Known limitation | Slower to incorporate new structures due to manual effort | Automated steps can produce domain definitions that differ from SCOP |

Both databases serve as gold standards for benchmarking protein structure comparison methods and for training machine learning approaches to structure classification. However, the two hierarchies result from different protocols, and the same protein can receive different classifications in each system. Ignoring these differences creates problems when the databases are used to train or benchmark automatic structure classification methods. A consistent mapping between SCOP and CATH can reduce errors made by structure comparison methods such as TM-Align and supports further applications in machine learning for protein structure classification.

## Understanding Protein Structure Classification

Protein structure classification organizes the tens of thousands of experimentally determined structures in the Protein Data Bank into a manageable framework. Web-based protein structure databases come in a wide variety of types and levels of information content. The atlases that describe each experimentally determined protein structure provide useful links, analyses, and schematic diagrams relating to its three-dimensional structure and biological function. The databases that classify three-dimensional structures by their folds are of great interest because they can reveal evolutionary relationships that may be hard to detect from sequence comparison alone.

The need for classification arises from a fundamental observation about protein evolution. Protein structures are often better conserved than protein sequences. When sequence similarity is too low for application of sequence-based homology search or phylogenetic methods, comparison of protein structures may provide an alternative means of uncovering deep evolutionary signal. Major protein structure databases such as SCOP and CATH hierarchically group protein structures, but they do not describe the specific evolutionary relationships within a hierarchical level. Structural phylogenies have the potential to fill this gap.

For researchers working with newly solved structures, especially those of unknown function, fold comparison servers and classification databases are particularly useful. The choice between SCOP and CATH affects how you interpret your structure and how you benchmark any computational methods you develop.

## Classification Hierarchies and Domain Definitions

### SCOP Hierarchy

SCOP organizes protein structures into a hierarchy based on evolutionary relationships and structural principles. The main levels are Class, Fold, Superfamily, and Family. The Class level describes the secondary structure composition of the protein, such as all-alpha, all-beta, or alpha/beta. The Fold level groups proteins that share the same arrangement of secondary structure elements in three-dimensional space. The Superfamily level brings together proteins that are likely to share a common evolutionary ancestor based on structural and functional evidence. The Family level contains proteins with clear sequence similarity and close evolutionary relationships.

The manual curation in SCOP means that each classification decision is made by expert inspection of the structure. This approach produces classifications that reflect careful biological judgment about evolutionary relationships. The tradeoff is that manual curation takes time, so SCOP releases updates less frequently than automated systems.

### CATH Hierarchy

CATH uses a different hierarchy with the levels Class, Architecture, Topology, and Homologous superfamily. The Class level, like SCOP, describes secondary structure content. The Architecture level describes the orientation of secondary structure elements without considering the connectivity between them. The Topology level, also called the fold level, groups proteins with the same arrangement and connectivity of secondary structure elements. The Homologous superfamily level groups proteins that share a common ancestor, identified through sequence, structural, and functional evidence.

CATH combines automated methods with manual curation. The automated steps allow CATH to process new structures more quickly than purely manual systems. The manual steps provide quality control and biological interpretation at key points in the classification process.

### Domain Definition Differences

One of the most significant practical differences between SCOP and CATH lies in how they define protein domains. A domain is a region of a protein that can fold independently and often has its own function. The two databases do not always agree on where domain boundaries lie within a multi-domain protein. Large and unexpected differences exist between SCOP and CATH with respect to their domain definitions as well as their hierarchic partitioning of the fold space on every level of the two classifications.

This disagreement matters for practical research. If you are comparing structures with a tool like TM-Align, the domain definitions you use affect the alignment results. A consistent mapping between SCOP and CATH can reduce errors made by structure comparison methods. Researchers have created benchmark sets and interactive browsers to support this mapping, which are useful for automated structure comparison and classification.

## Curation Methods and Update Frequency

### Manual Curation in SCOP

SCOP represents a purely manual approach to protein structure classification. Expert curators inspect each structure and make classification decisions based on their understanding of protein evolution, structure, and function. This approach produces high-quality classifications that reflect detailed biological knowledge. The manual process also means that SCOP classifications can capture subtle evolutionary relationships that automated methods might miss.

The limitation of manual curation is scalability. The number of protein structures deposited in the Protein Data Bank continues to grow, and manual classification cannot keep pace with the rapid generation of new structures. This results in a bias toward unannotated or unclassified structures in the database.

### Automated and Manual Combination in CATH

CATH uses a combination of automated and manual methods. Automated steps process new structures quickly and assign preliminary classifications. Manual curation provides quality control and refines classifications where automated methods are uncertain. This hybrid approach allows CATH to process more structures than purely manual systems while maintaining a level of expert oversight.

The automated components of CATH also support integration with other bioinformatics resources. CATH-Gene3D is one of the protein signature databases that contribute to InterPro, a combined annotation tool that integrates protein signatures from multiple databases. InterPro provides structural information from the Protein Data Bank, its classification in CATH and SCOP, as well as homology models from ModBase. This integration makes CATH classifications accessible through the InterPro resource for automatic annotation of proteins.

### Update Frequency Considerations

The update frequency of a classification database affects how current your analysis will be. If you are working with a newly solved structure, you may need to wait for the next database release before your structure appears in the classification. Manual curation in SCOP means longer intervals between releases. The automated components of CATH allow more frequent incorporation of new structures.

For research questions that depend on the most current classification of recently deposited structures, the update frequency of the database is a practical consideration. For research questions that depend on careful evolutionary analysis, the depth of manual curation may matter more than update speed.

## Coverage and Integration with Other Resources

### Protein Data Bank Coverage

Both SCOP and CATH aim to classify the structures available in the Protein Data Bank. The Protein Data Bank contains a large and growing number of experimentally determined protein structures. As of December 2019, the Protein Data Bank contained in excess of 158,000 entries. The growth of the database creates ongoing challenges for classification efforts.

Approximately two-thirds of the protein chains in SCOP, CATH, and FSSP are common to all three databases. Despite employing different methods and basing their systems on different rules of protein structure and taxonomy, SCOP, CATH, and FSSP agree on the majority of their classifications. Discrepancies and inconsistencies are accounted for by a small number of explanations.

### InterPro Integration

CATH classifications are integrated into InterPro through the CATH-Gene3D signature database. InterPro combines protein signatures from eleven databases, including CATH-Gene3D, HAMAP, PANTHER, Pfam, PIRSF, PRINTS, ProDom, PROSITE, SMART, SUPERFAMILY, and TIGRFAMs. These databases use approaches ranging from characterising small conserved motifs to using hidden Markov models that describe the conservation of residues over entire domains or whole proteins.

InterPro is an open-source protein resource used for the automatic annotation of proteins. It is scalable to the analysis of entire new genomes through the use of a downloadable version of InterProScan, which can be incorporated into an existing local pipeline. If your research involves genome annotation or large-scale protein analysis, the InterPro integration may make CATH classifications more accessible than SCOP classifications.

### SUPERFAMILY and SCOP

SCOP classifications are also made available through derived resources. The SUPERFAMILY database uses hidden Markov models to assign SCOP classifications to protein sequences. This allows researchers to search sequence databases for proteins that are likely to adopt structures classified in SCOP. If you have a sequence without a known structure, SUPERFAMILY can help you predict its structural classification based on SCOP.

## Practical Workflow for Database Selection

### Step 1: Define Your Research Question

The first decision point is whether your research question requires evolutionary relationships, structural similarity, or both. If you need to understand deep evolutionary relationships among proteins where sequence similarity is too low for sequence-based methods, structural comparison is your primary tool. Both SCOP and CATH can help, but you need to understand what each database does and does not provide.

SCOP and CATH hierarchically group protein structures, but they do not describe the specific evolutionary relationships within a hierarchical level. If you need to reconstruct evolutionary relationships within a superfamily, you will need to perform structural phylogenetics yourself. Structural phylogenies have the potential to fill the gap left by the hierarchical classifications.

### Step 2: Assess Your Structure Set

Consider the characteristics of the structures you are working with. Structural phylogenetics is best employed where structures have very similar lengths. Shape fluctuations generated during molecular dynamics simulations impact pairwise comparisons, but not so drastically as to eliminate evolutionary signal. If your structures vary widely in length, you should account for this in your analysis.

If you are working with a newly solved structure of unknown function, fold comparison servers and classification databases are particularly useful. The choice between SCOP and CATH may depend on which database has already classified structures similar to yours.

### Step 3: Evaluate Benchmarking Requirements

If you are developing or testing structure comparison methods, the choice of classification database affects your benchmark. SCOP and CATH are widely used as gold standards to benchmark novel protein structure comparison methods as well as to train machine learning approaches for protein structure classification and prediction. The two hierarchies result from different protocols, which may result in differing classifications of the same protein.

Ignoring these differences leads to problems when the databases are used to train or benchmark automatic structure classification methods. A consistent mapping between SCOP and CATH defines a consistent benchmark set that largely reduces errors made by structure comparison methods such as TM-Align. If you are benchmarking, consider using a consensus approach that accounts for agreements between databases.

### Step 4: Consider Update Needs

Determine how current your classifications need to be. If you are analyzing structures deposited in the last year, check whether the database release you plan to use includes those structures. Manual curation in SCOP means that new structures may not appear in SCOP for some time after deposition. The automated components of CATH may incorporate new structures more quickly.

### Step 5: Check Integration Requirements

Consider whether you need to integrate your classifications with other bioinformatics resources. If you are annotating genomes or analyzing large sequence datasets, the InterPro integration of CATH classifications may be more convenient. If you are working with sequences and need to predict structural classifications, the SUPERFAMILY resource based on SCOP may be more appropriate.

## Options and Tradeoffs in Database Selection

### When to Use SCOP

SCOP is well suited for research questions that require careful manual curation and detailed evolutionary analysis. The expert judgment applied in SCOP classifications can capture evolutionary relationships that automated methods might miss. If you are studying a protein family where subtle structural differences distinguish evolutionary groups, SCOP classifications may provide the resolution you need.

SCOP is also appropriate when you need to understand the biological reasoning behind a classification. The manual curation process means that each classification reflects explicit biological judgment. This can be valuable when you are interpreting the evolutionary history of a protein family.

The limitation of SCOP is its update frequency. If you need classifications for recently deposited structures, you may need to wait for the next SCOP release. The manual curation process cannot keep pace with the rapid generation of new structures.

### When to Use CATH

CATH is well suited for large-scale analyses where automated processing is an advantage. The combination of automated and manual methods allows CATH to process more structures than purely manual systems. If you are analyzing a large set of structures or annotating a genome, CATH classifications may be more complete and current.

CATH is also appropriate when you need to integrate your analysis with other bioinformatics resources. The CATH-Gene3D integration with InterPro makes CATH classifications accessible through a widely used annotation tool. If your workflow already uses InterPro, CATH classifications are readily available.

The limitation of CATH is that automated steps can produce domain definitions that differ from SCOP. If your analysis depends on specific domain boundaries, you should check whether CATH and SCOP agree on the domains in your structures.

### When to Use Both

For many research questions, using both SCOP and CATH is the best approach. The two databases provide orthogonal features that can be exploited for automated structure comparison and classification. A consistent mapping of SCOP and CATH can be exploited for automated structure comparison and classification.

Using both databases allows you to identify cases where the two systems agree, which increases confidence in the classification. Cases where the two systems disagree highlight regions of the protein fold space where classification is uncertain. These disagreements can be biologically informative, revealing proteins that sit at the boundaries between folds or superfamilies.

### When to Use FSSP or Other Resources

FSSP represents a purely automated approach to protein structure classification. The systematic comparison of SCOP, CATH, and FSSP shows that approximately two-thirds of the protein chains in each database are common to all three databases. Despite employing different methods, the databases agree on the majority of their classifications.

FSSP may be appropriate when you need fully automated classifications that can be regenerated quickly as new structures are deposited. The automated approach means that FSSP classifications are not subject to the delays of manual curation. However, the lack of manual curation means that FSSP classifications may not capture the biological nuances that expert curators identify.

## Observations and Measurements for Database Comparison

### Measuring Classification Agreement

When comparing SCOP and CATH classifications, you can measure the agreement between the two systems at each hierarchical level. The two hierarchies result from different protocols, which may result in differing classifications of the same protein. Large and unexpected differences exist between SCOP and CATH with respect to their domain definitions as well as their hierarchic partitioning of the fold space on every level of the two classifications.

For practical research, you should measure the agreement between SCOP and CATH for your specific structure set. This measurement tells you how much confidence you can place in classifications that come from a single database. If SCOP and CATH agree on your structures, you can be more confident in the classification. If they disagree, you need to investigate the source of the disagreement.

### Recording Domain Boundary Differences

Domain boundary differences between SCOP and CATH are a common source of classification disagreement. When you record classifications for your structures, note which database defined the domain boundaries. If you are comparing structures across databases, use consistent domain definitions.

A consistent mapping between SCOP and CATH can reduce errors made by structure comparison methods such as TM-Align. If you are benchmarking structure comparison tools, use a benchmark set that accounts for the differences between SCOP and CATH domain definitions.

### Tracking Update Lags

Record the release dates of the database versions you use in your analysis. If you are comparing classifications across time, note when each structure was added to each database. This information helps you interpret differences that arise from update lags instead of from genuine classification disagreements.

## Records and Documentation for Reproducible Analysis

### Documenting Database Versions

Reproducible analysis requires documentation of the exact database versions you used. Record the release number or date for both SCOP and CATH. If you use derived resources such as SUPERFAMILY or InterPro, record the version of those resources as well.

The importance of version documentation is well established in bioinformatics training and practice. Resources such as the [Galaxy Training Network](https://training.galaxyproject.org/) emphasize accessible workflow training and reproducibility context. The [nf-core documentation](https://nf-co.re/docs) describes community pipeline standards for reproducible workflows. The [Carpentries lessons](https://carpentries.org/lessons) on data and computing fundamentals also stress the importance of reproducible workflows.

### Recording Classification Parameters

If you use automated classification tools, record the parameters you used. Different parameter settings can produce different classifications. For structure comparison tools, record the alignment method and any thresholds you applied.

### Maintaining Analysis Logs

Keep a log of your classification analyses, including the input structures, the database versions, and the results. This log supports reproducibility and helps you diagnose problems if classifications change in later database releases.

## Quality Controls for Classification Analysis

### Checking Domain Assignment Consistency

Before relying on classifications from either SCOP or CATH, check whether the domain assignments for your structures are consistent across databases. If SCOP and CATH assign different domain boundaries to the same structure, investigate the source of the difference. The difference may arise from genuine structural ambiguity or from different classification protocols.

### Validating with Structural Alignment

Validate classifications by performing structural alignments of your structures against representatives from the assigned fold or superfamily. Structural alignment provides independent evidence for or against the classification. If your structure aligns well with members of the assigned class, the classification is supported. If the alignment is poor, the classification may be incorrect.

### Cross-Checking with Sequence Resources

Cross-check structural classifications with sequence-based resources. Protein signature databases describe protein families, functional domains, or conserved sites within related groups of proteins. If sequence-based signatures support the structural classification, you can be more confident in the result.

### Monitoring for Classification Changes

Monitor whether classifications change between database releases. If a structure moves from one fold to another in a new release, investigate the reason for the change. Classification changes can result from improved methods, new structural data, or corrections to previous classifications.

## Common Failure Patterns in Database Selection

### Assuming SCOP and CATH Are Interchangeable

A common failure is assuming that SCOP and CATH classifications can be used interchangeably. The two hierarchies result from different protocols, which may result in differing classifications of the same protein. Ignoring these differences leads to problems when the databases are used to train or benchmark automatic structure classification methods.

### Using Outdated Database Versions

Another failure pattern is using outdated database versions without documenting the version. Classifications change between releases, and analyses based on outdated versions may not reflect current knowledge. Always document the database version you used.

### Ignoring Domain Definition Differences

Researchers who ignore domain definition differences between SCOP and CATH can produce incorrect results. If you compare structures using domain boundaries from one database while classifications come from another, your results may be inconsistent. Use consistent domain definitions throughout your analysis.

### Overinterpreting Hierarchical Relationships

A failure pattern in evolutionary analysis is overinterpreting the hierarchical relationships in SCOP or CATH as direct evolutionary statements. SCOP and CATH hierarchically group protein structures, but they do not describe the specific evolutionary relationships within a hierarchical level. If you need evolutionary relationships within a superfamily, you must perform structural phylogenetics.

### Neglecting Update Frequency

Researchers who need current classifications may fail to check whether the database release includes their structures. Manual curation in SCOP means that new structures may not appear for some time after deposition. Check the release date and coverage of the database version you plan to use.

## Limitations of SCOP and CATH

### Manual Curation Bottlenecks

The manual curation in SCOP creates a bottleneck in classification throughput. The ability to manually classify and annotate sequences cannot keep pace with their rapid generation, resulting in an increased bias toward unannotated sequence. This limitation affects the coverage and currency of SCOP classifications.

### Automated Method Uncertainties

The automated components of CATH can produce classifications that differ from manual classifications. The automated steps may not capture subtle biological features that expert curators identify. These differences are most apparent in domain definitions and in the partitioning of the fold space.

### Structural Phylogeny Challenges

Structural phylogenetics, which can complement the hierarchical classifications, has its own limitations. Structural phylogenies are best employed where structures have very similar lengths. Shape fluctuations generated during molecular dynamics simulations impact pairwise comparisons, but not so drastically as to eliminate evolutionary signal. Researchers using structural phylogenetics need methods for assessing confidence in their trees.

### Coverage Gaps

Both SCOP and CATH have coverage gaps. Not all structures in the Protein Data Bank are classified in either database. The growth of the Protein Data Bank continues to challenge classification efforts. If your structure is not classified, you may need to use fold comparison servers to identify related structures.

## Safety and Reproducibility Context

### Reproducibility Standards

Reproducibility is a core concern in bioinformatics analysis. Training resources from the [Galaxy Training Network](https://training.galaxyproject.org/) emphasize accessible workflow training and reproducibility context. The [nf-core documentation](https://nf-co.re/docs) describes community pipeline standards for reproducible workflows. The [Carpentries lessons](https://carpentries.org/lessons) provide foundational computing and data training that supports reproducible analysis.

When you use SCOP or CATH classifications in your research, follow reproducibility standards by documenting your database versions, parameters, and analysis steps. This documentation allows others to reproduce your results and to understand the basis for your classifications.

### Data Management Practices

Good data management practices support reliable classification analysis. Store your input structures, classification results, and analysis scripts in organized locations. Document the provenance of your data, including the database versions and retrieval dates.

### Professional Escalation Criteria

Some classification questions require professional judgment beyond what database documentation can provide. Consider escalating to a structural bioinformatics specialist or a protein evolution expert when you encounter any of the following situations:

- SCOP and CATH disagree on the classification of a structure that is central to your research question
- Your structure does not fit cleanly into any existing fold or superfamily
- You need to make evolutionary claims based on structural classifications and you are not trained in phylogenetic methods
- You are developing a benchmark set for structure comparison methods and need guidance on handling classification disagreements
- You are annotating a genome and need to interpret conflicting structural and sequence-based classifications

## Building a Decision Matrix for SCOP and CATH Selection

A practical decision matrix helps you move from abstract database differences to a concrete selection for your specific research context. The matrix below translates the key operational differences between SCOP and CATH into scored criteria that you can apply to your own project. This approach is particularly useful when you are working with a mixed structure set or when your research question spans multiple classification needs.

### Decision Matrix Criteria

The decision matrix uses five criteria that map directly to the documented differences between SCOP and CATH. Each criterion receives a weight based on your research priorities, and each database receives a score based on its documented performance for that criterion.

| Criterion | SCOP Score | CATH Score | Weight for Your Project | Weighted SCOP | Weighted CATH |
| --- | --- | --- | --- | --- | --- |
| Curation depth for evolutionary analysis | 5 | 3 | Assign 1 to 5 | Multiply | Multiply |
| Update frequency for new structures | 2 | 4 | Assign 1 to 5 | Multiply | Multiply |
| Integration with sequence annotation tools | 3 | 5 | Assign 1 to 5 | Multiply | Multiply |
| Domain boundary consistency for benchmarking | 4 | 3 | Assign 1 to 5 | Multiply | Multiply |
| Coverage of your specific structure set | Check per structure | Check per structure | Assign 1 to 5 | Multiply | Multiply |

The scores in the table reflect the documented characteristics of each database. SCOP scores higher on curation depth because it uses purely manual classification by expert inspection. CATH scores higher on update frequency because it combines automated methods with manual curation, allowing more rapid incorporation of new structures. CATH scores higher on integration because CATH-Gene3D contributes to InterPro, which combines protein signatures from eleven databases including CATH-Gene3D, HAMAP, PANTHER, Pfam, PIRSF, PRINTS, ProDom, PROSITE, SMART, SUPERFAMILY, and TIGRFAMs.

### Applying the Matrix to Common Research Scenarios

For a benchmarking project where you are testing a new structure comparison method, the domain boundary consistency criterion should receive the highest weight. SCOP and CATH are widely used as gold standards to benchmark novel protein structure comparison methods as well as to train machine learning approaches for protein structure classification and prediction. The two hierarchies result from different protocols, which may result in differing classifications of the same protein. Ignoring such differences leads to problems when being used to train or benchmark automatic structure classification methods. In this scenario, you should assign a weight of 5 to domain boundary consistency and a weight of 3 or lower to update frequency.

For a genome annotation project where you need to classify thousands of predicted protein sequences, the integration criterion should dominate your decision. InterPro is an open-source protein resource used for the automatic annotation of proteins, and is scalable to the analysis of entire new genomes through the use of a downloadable version of InterProScan, which can be incorporated into an existing local pipeline. CATH classifications are accessible through this pipeline, making CATH the practical choice for large-scale annotation work.

For a deep evolutionary study where you are investigating relationships among proteins with very low sequence similarity, the curation depth criterion matters most. Protein structures are often better conserved than sequences, and comparison of protein structures may provide an alternative means of uncovering deep evolutionary signal. The manual curation in SCOP can capture subtle evolutionary relationships that automated methods might miss. However, you should also note that SCOP and CATH hierarchically group protein structures, but they do not describe the specific evolutionary relationships within a hierarchical level. You will still need to perform structural phylogenetics yourself.

### Recording Your Matrix Scores

Document your decision matrix in your analysis log. Record the weights you assigned to each criterion, the scores for each database, and the final weighted totals. This documentation supports reproducibility and helps you revisit your decision if your research question evolves.

The importance of documenting analysis decisions is emphasized across bioinformatics training resources. The [Galaxy Training Network](https://training.galaxyproject.org/) provides accessible workflow training and reproducibility context. The [nf-core documentation](https://nf-co.re/docs) describes community pipeline standards for reproducible workflows. The [Carpentries lessons](https://carpentries.org/lessons) provide foundational computing and data training that supports reproducible analysis practices.

### Troubleshooting the Matrix Results

If your decision matrix produces a close score between SCOP and CATH, investigate the specific structures in your set that drive the uncertainty. Check whether SCOP and CATH agree on the classification of those structures. Approximately two-thirds of the protein chains in SCOP, CATH, and FSSP are common to all three databases, and despite employing different methods, the databases agree on the majority of their classifications. If your structures fall in the agreeing majority, the choice between databases matters less. If your structures fall in the disagreeing minority, you need to understand the source of the disagreement.

Large and unexpected differences exist between SCOP and CATH with respect to their domain definitions as well as their hierarchic partitioning of the fold space on every level of the two classifications. When your matrix produces a close result, examine the domain definitions for your specific structures. Domain boundary differences are a common source of classification disagreement and can affect downstream analysis.

### Building a Consensus Classification Set

When your research requires high-confidence classifications, build a consensus set that includes only structures where SCOP and CATH agree. A consistent mapping between SCOP and CATH defines a consistent benchmark set which is shown to largely reduce errors made by structure comparison methods such as TM-Align. This consensus approach is particularly valuable for benchmarking and for training machine learning methods.

To build a consensus set, retrieve classifications for your structures from both databases. Compare the assignments at each hierarchical level. Record which structures receive the same classification in both systems and which receive different classifications. Use the agreeing structures for your high-confidence analyses and investigate the disagreeing structures separately.

The consensus approach has useful further applications beyond benchmarking. It can support machine learning methods being trained for protein structure classification. It also allows you to extract additional connections in the topology of the protein fold space from the orthogonal features contained in SCOP and CATH.

### Handling Structures Not Yet Classified

Both SCOP and CATH have coverage gaps. The manual classification process cannot keep pace with the rapid generation of new structures, resulting in an increased bias toward unannotated sequence. If your structure is not yet classified in either database, you have several options.

First, check whether a more recent release of either database includes your structure. The automated components of CATH may incorporate new structures more quickly than the manual curation in SCOP. Second, use fold comparison servers to identify structures related to yours. These servers are particularly useful for newly solved structures, and especially those of unknown function. Third, consider using sequence-based prediction through SUPERFAMILY, which uses hidden Markov models to assign SCOP classifications to protein sequences.

### Validating Your Database Choice

After you select a database using the decision matrix, validate your choice with a small pilot analysis. Retrieve classifications for a sample of your structures from the selected database. Perform structural alignments of your structures against representatives from the assigned fold or superfamily. If the alignments support the classifications, your database choice is validated. If the alignments are poor, reconsider your selection.

Cross-check your structural classifications with sequence-based resources. Protein signature databases describe protein families, functional domains, or conserved sites within related groups of proteins. If sequence-based signatures support the structural classification, you can be more confident in the result. The InterPro database combines protein signatures from eleven databases to increase their value as protein classification tools.

### Monitoring Classification Stability

After you commit to a database for your project, monitor whether classifications change between releases. If a structure moves from one fold to another in a new release, investigate the reason for the change. Classification changes can result from improved methods, new structural data, or corrections to previous classifications.

Record the release dates of the database versions you use in your analysis. If you are comparing classifications across time, note when each structure was added to each database. This information helps you interpret differences that arise from update lags instead of from genuine classification disagreements.

### Escalating to Professional Judgment

Some classification decisions require professional judgment beyond what the decision matrix can provide. Consider escalating to a structural bioinformatics specialist or a protein evolution expert when you encounter any of the following situations:

- Your decision matrix produces conflicting results across multiple criteria and the conflict affects a central part of your research
- SCOP and CATH disagree on the classification of a structure that is central to your research question
- Your structure does not fit cleanly into any existing fold or superfamily
- You need to make evolutionary claims based on structural classifications and you are not trained in phylogenetic methods
- You are developing a benchmark set for structure comparison methods and need guidance on handling classification disagreements
- You are annotating a genome and need to interpret conflicting structural and sequence-based classifications

### Integrating the Decision Matrix with Existing Workflows

The decision matrix integrates with standard bioinformatics workflows. If you use workflow platforms for your analysis, incorporate the database selection step as an explicit stage in your pipeline. The [Galaxy Training Network](https://training.galaxyproject.org/) provides tutorials for building reproducible analysis workflows. The [nf-core documentation](https://nf-co.re/docs) describes community pipeline standards that support reproducible workflow configuration.

If you use R or Bioconductor for your downstream analysis, record your database selection and classification results in a structured format that supports reproducible analysis. The [Bioconductor Project](https://bioconductor.org/) provides official package, workflow, installation, and reproducible genomic-analysis documentation. Storing your decision matrix and classification results in a structured format allows you to revisit your decisions and to share your methods with collaborators.

### Common Mistakes in Applying the Decision Matrix

A common mistake is assigning equal weights to all criteria without considering the specific demands of your research question. The decision matrix is only useful if the weights reflect your actual priorities. A benchmarking project and a genome annotation project should produce different database selections because they have different priorities.

Another mistake is failing to update the matrix when your research question changes. If you start with a benchmarking project and later expand to include evolutionary analysis, revisit your weights. The database that was optimal for benchmarking may not be optimal for evolutionary analysis.

A third mistake is ignoring the domain definition differences when interpreting your matrix results. Even if your matrix selects one database, you should check whether the domain definitions in that database match your analytical needs. If you are comparing structures using domain boundaries from one database while classifications come from another, your results may be inconsistent.

### Using the Matrix for Team Decisions

When multiple researchers are involved in a project, use the decision matrix as a structured discussion tool. Have each researcher assign weights independently, then compare the results. Disagreements about weights often reveal different assumptions about the research goals. Resolving these disagreements before you commit to a database saves time and prevents rework later.

Document the final weights and the rationale for each weight assignment. This documentation is particularly valuable for multi-year projects where team members may change. New team members can understand why a particular database was selected and can revisit the decision if the research direction changes.

### Reviewing the Matrix Against New Database Releases

Revisit your decision matrix when new releases of SCOP or CATH become available. A release that significantly expands coverage or improves domain definitions could change your database selection. The automated components of CATH may incorporate new structures more quickly, but a major SCOP release could provide improved classifications for your specific structure set.

Check the release notes for both databases to understand what changed. If the changes affect your structures or your research question, update your matrix scores and weights accordingly. Record the date of your matrix review and the database versions you considered.

## Frequently Asked Questions

### What is the main difference between SCOP and CATH?

The main difference is the classification approach. SCOP uses purely manual curation by expert inspection, while CATH combines automated methods with manual curation. This difference affects domain definitions, update frequency, and the partitioning of the fold space. The two hierarchies result from different protocols, which may result in differing classifications of the same protein.

### Which database should I use for benchmarking structure comparison methods?

For benchmarking, you should use a consistent benchmark set that accounts for the differences between SCOP and CATH. A consistent mapping between SCOP and CATH defines a benchmark set that largely reduces errors made by structure comparison methods such as TM-Align. Using both databases and focusing on cases where they agree can improve benchmark reliability.

### How do SCOP and CATH define protein domains differently?

SCOP and CATH use different protocols for defining domain boundaries within multi-domain proteins. Large and unexpected differences exist between SCOP and CATH with respect to their domain definitions as well as their hierarchic partitioning of the fold space on every level of the two classifications. These differences affect which regions of a protein are classified as separate domains.

### Can I use SCOP and CATH to study evolutionary relationships?

SCOP and CATH hierarchically group protein structures, but they do not describe the specific evolutionary relationships within a hierarchical level. For deep evolutionary relationships where sequence similarity is too low for sequence-based methods, structural comparison provides an alternative. Structural phylogenies have the potential to fill the gap left by the hierarchical classifications.

### How often are SCOP and CATH updated?

SCOP uses manual curation, which takes time and results in less frequent updates. CATH combines automated and manual methods, allowing more frequent incorporation of new structures. The manual classification process cannot keep pace with the rapid generation of new structures, resulting in a bias toward unannotated sequence.

### What is the relationship between CATH and InterPro?

CATH classifications are integrated into InterPro through the CATH-Gene3D signature database. InterPro combines protein signatures from eleven databases and provides structural information from the Protein Data Bank, its classification in CATH and SCOP, as well as homology models from ModBase. InterPro is scalable to the analysis of entire new genomes through InterProScan.

### How can I assign SCOP classifications to protein sequences?

The SUPERFAMILY database uses hidden Markov models to assign SCOP classifications to protein sequences. This allows researchers to search sequence databases for proteins that are likely to adopt structures classified in SCOP. This approach is useful when you have sequences without known structures.

### What should I do if SCOP and CATH disagree on my structure?

If SCOP and CATH disagree on the classification of your structure, investigate the source of the disagreement. The disagreement may arise from different domain definitions or from different partitioning of the fold space. A consistent mapping between SCOP and CATH can help you understand the relationship between the two classifications. If the disagreement affects a central part of your research, consider consulting a structural bioinformatics specialist.

## Related Bioinformatics Guides

- [Structural Comparison and Alignment Algorithms for Protein 3D Structures](/knowledge/bioinformatics/structural-comparison-and-alignment-algorithms-for-protein-3d-structures)
- [RNA-Seq vs qPCR: Validation and Comparison](/knowledge/bioinformatics/rna-seq-vs-qpcr-validation-and-comparison)
- [Metagenomic Binning Tools Benchmark: How to Evaluate and Choose](/knowledge/bioinformatics/metagenomic-binning-tools-benchmark-how-to-evaluate-and-choose)
- [Single-Cell Sequencing Services: How to Choose a Provider](/knowledge/bioinformatics/single-cell-sequencing-services-how-to-choose-a-provider)
- [Metagenomics vs Metabarcoding: Choosing the Right Approach for Your Study](/knowledge/bioinformatics/metagenomics-vs-metabarcoding-choosing-the-right-approach-for-your-study)

## References and Further Reading

- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Bioconductor](https://bioconductor.org/). Bioconductor Project.
- [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.
- [nf-core Documentation](https://nf-co.re/docs). nf-core.
- [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.
- [Structural Phylogenetics with Confidence.](https://pubmed.ncbi.nlm.nih.gov/32302382). Molecular biology and evolution, 2020.
- [InterPro protein classification.](https://pubmed.ncbi.nlm.nih.gov/21082426). Methods in molecular biology (Clifton, N.J.), 2011.
- [A systematic comparison of protein structure classifications: SCOP, CATH and FSSP.](https://pubmed.ncbi.nlm.nih.gov/10508779). Structure (London, England : 1993), 1999.
- [Systematic comparison of SCOP and CATH: a new gold standard for protein structure analysis.](https://pubmed.ncbi.nlm.nih.gov/19374763). BMC structural biology, 2009.
- [Protein Structure Databases.](https://pubmed.ncbi.nlm.nih.gov/27115626). Methods in molecular biology (Clifton, N.J.), 2016.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.