# How to Retrieve and Interpret Protein Annotations from UniProtKB: A Step-by-Step Guide for Proteomics Researchers

Proteomics researchers who identify proteins through mass spectrometry or other high-throughput methods face a common problem: the protein list from a search engine contains accession numbers, but the biological meaning of those accessions requires functional annotation. UniProtKB is the central resource for this task, providing protein sequences with functional information that is freely accessible to all users. This article walks through the practical steps of retrieving annotations from UniProtKB, interpreting the key fields that matter for proteomics follow-up work, and exporting the data in formats suitable for downstream analysis. The workflow assumes you have a list of protein accessions from a typical proteomics experiment and need to map those accessions to functional information such as gene ontology terms, subcellular locations, post-translational modifications, and expression patterns.

## Understanding the UniProtKB Database Structure

UniProtKB is maintained by the UniProt consortium, which brings together the Swiss Institute of Bioinformatics, the European Bioinformatics Institute, and the Protein Information Resource. The database exists to provide the scientific community with a central resource for protein sequences and functional information. The knowledgebase is split into two sections that differ fundamentally in how their entries are created and maintained.

### UniProtKB/Swiss-Prot and UniProtKB/TrEMBL

UniProtKB/Swiss-Prot contains manually annotated entries. Biologists read the scientific literature and curate functional information for each protein in this section. These entries include cross-links to roughly 100 external databases, giving users access to additional information and tools beyond what UniProtKB itself stores. The manual curation process means Swiss-Prot entries carry a higher confidence level for their annotations.

UniProtKB/TrEMBL contains computer-annotated entries. These are produced through automated pipelines that predict functional information for proteins that have not been studied experimentally. The automated annotation methods used by UniProtKB predict information for unreviewed entries describing unstudied proteins. The distinction matters for proteomics researchers because the review status of an entry affects how much trust you can place in the functional claims.

The production pipeline has been modified to limit the sequences available in UniProtKB to high-quality, non-redundant reference proteomes. This change means the database is increasingly focused on well-supported protein sequences instead of including every possible translation product from every genome. For proteomics researchers, this reduces the noise in accession mapping and improves the likelihood that your identified peptides correspond to curated entries.

### Supporting Databases: UniRef and UniParc

Beyond UniProtKB itself, the UniProt consortium maintains supplementary databases that serve different purposes. UniRef defines clusters of protein sequences that share 100, 90, or 50 percent identity. These clusters help researchers work with redundant sequence sets by grouping similar proteins. UniParc stores and maps all publicly available protein sequence data, including obsolete data that has been excluded from UniProtKB. If you encounter an accession that no longer resolves in UniProtKB, UniParc is the place to check whether that sequence still exists in the archival record.

The UniProt website receives about 800,000 unique visitors per month and provides 10 searchable datasets and four main tools. The key datasets are UniProtKB, UniRef, UniParc, and Proteomes, which contains protein sets for completely sequenced genomes. Other supporting datasets include literature citations, taxonomy, and subcellular locations.

## At a Glance: Annotation Retrieval Decision Table

| Scenario | Recommended Approach | Key Fields to Extract | Export Format |
| --- | --- | --- | --- |
| Small protein list under 25 accessions from a targeted proteomics experiment | Manual entry lookup through the UniProtKB search interface | Function, subcellular location, PTM, expression, sequence | Tab-separated values or FASTA for sequence records |
| Medium list of 25 to 500 accessions from a discovery proteomics run | Batch retrieval using the Retrieve/ID mapping tool | Gene ontology, reviewed status, protein existence, cross-references | Tab-separated values with custom columns |
| Large list over 500 accessions or repeated analyses across multiple experiments | Programmatic access through the REST API or SPARQL endpoint | All annotation fields with evidence codes and source attribution | JSON or XML for automated parsing |
| Accessions that fail to map or return unexpected proteins | Verify accession format, check UniParc for obsolete sequences, confirm taxonomy | Review status, sequence version, entry history | Record unmapped accessions in your laboratory notebook |

## Preparing Your Accession List for Retrieval

Before you begin any retrieval workflow, you need to know what type of identifiers you have. Proteomics search engines typically report UniProtKB accessions, but some tools report RefSeq identifiers, Ensembl protein identifiers, or gene symbols. The UniProtKB Retrieve/ID mapping tool accepts multiple identifier types, but the success rate depends on the format you provide.

### Identifier Formats and Their Implications

UniProtKB accessions follow a specific pattern. A typical accession is six to ten characters long, starting with a letter. Swiss-Prot accessions and TrEMBL accessions both follow this pattern, but the entry name and gene name fields differ. If your search engine reports accessions with a species-specific prefix or a version suffix, you may need to strip those before submitting the list.

The NCBI maintains its own set of databases that serve complementary roles in the bioinformatics ecosystem. The NCBI data resources include sequence databases, search systems, and analysis services that can be used to cross-check protein identifiers when UniProtKB mapping fails. For example, if a peptide maps to a RefSeq protein that has no UniProtKB counterpart, you can use the NCBI resources to examine the sequence record directly.

### Cleaning Your Input List

Remove duplicate accessions before submission. Duplicates waste processing time and can confuse downstream counting if you use the annotation file to quantify protein groups. Check for whitespace characters, especially if you copied the list from a spreadsheet or a PDF. Convert the list to plain text with one accession per line. Verify that the accessions match the taxonomy of your study organism. A common error is submitting human accessions when the experiment used mouse tissue, which produces either failed mappings or incorrect annotations.

## Step-by-Step Retrieval Through the UniProtKB Website

The UniProtKB website is the primary means to access UniProt data. The search interface supports both simple queries and advanced search construction. For proteomics researchers, the most efficient path depends on the size of your accession list.

### Single Protein Lookup

For a single protein, type the accession into the main search box on the UniProtKB homepage. The search will return the entry if the accession exists in the current release. The entry page displays the annotation sections in a structured layout. The Function section appears near the top and provides a summary of the protein's biological role. The Subcellular location section follows, describing where the protein is found within the cell. The PTM section lists post-translational modifications with their positions and the evidence supporting each modification.

The website includes a tab that links protein information to genomic information. This tab is useful when you need to connect your proteomics findings to genomic context, such as confirming that a protein isoform corresponds to a specific transcript variant.

### Batch Retrieval with Retrieve/ID Mapping

For lists of multiple accessions, use the Retrieve/ID Mapping tool. This tool accepts a pasted list or an uploaded file. Select the identifier type that matches your input, then select the destination database. For proteomics work, the most common destination is UniProtKB, but you can also map to Gene Ontology annotations or other resources.

After the mapping completes, you can select which columns to include in the output. The default output includes the accession, entry name, gene name, and organism. For functional annotation, add columns for Function, Subcellular location, PTM, Expression, and Gene ontology. The column selection interface allows you to build a custom output that matches your analysis needs.

### Advanced Search and Query Building

The advanced search interface lets you construct queries that combine multiple conditions. For example, you can search for all human proteins annotated with a specific subcellular location that also have a phosphorylation site at a particular position. The advanced search supports field-specific queries, and the query builder provides a visual interface for combining conditions.

The search functionality leverages the underlying data structure, including the Rhea knowledgebase for enzyme annotation. Rhea is a comprehensive expert-curated knowledgebase of biochemical reactions that describes reaction participants using the ChEBI ontology. The integration of Rhea into UniProtKB means that enzyme annotations are computationally tractable and searchable. If your proteomics experiment identifies enzymes, you can search for proteins by their reaction participants or by the ChEBI identifiers of their substrates and products.

## Interpreting Key Annotation Fields

The annotation fields in UniProtKB carry specific meanings that affect how you should interpret them in a proteomics context. Understanding these fields prevents common misinterpretations that lead to incorrect biological conclusions.

### Function Field

The Function field provides a textual summary of the protein's biological role. This text is curated from the scientific literature for Swiss-Prot entries and predicted by automated methods for TrEMBL entries. The text often includes information about catalytic activity, binding partners, and participation in biological processes. For enzymes, the catalytic activity is described using Rhea reactions, which provide a standardized representation of the biochemical reaction.

When interpreting the Function field, note whether the entry is reviewed or unreviewed. A reviewed entry has human-curated function information that has been checked against the literature. An unreviewed entry has predicted function information that may be incomplete or incorrect. The review status is displayed prominently on the entry page.

### Subcellular Location Field

The Subcellular location field describes where the protein is found in the cell. This information is critical for interpreting proteomics data because it helps you assess whether a protein identification makes biological sense. For example, a plasma membrane protein identified in a nuclear fraction may indicate contamination or a novel localization that warrants further investigation.

The subcellular location annotations include evidence codes that indicate the type of evidence supporting the location claim. Experimental evidence is stronger than predicted evidence. When you use subcellular location annotations to interpret your proteomics data, weight the evidence codes appropriately.

### Post-Translational Modification Field

The PTM field lists post-translational modifications with their positions and the evidence supporting each modification. The annotation of PTMs is an important task for UniProtKB curators, and the database uses text-mining tools to help curators track new publications on PTMs. The text-mining system extracts sentences from the scientific literature that contain information on the type and site of modification, and it provides a ranked list of protein candidates for the modification.

For proteomics researchers, the PTM field is particularly valuable because it provides a reference set of known modification sites. When your mass spectrometry analysis identifies a phosphorylation site, you can check whether that site is already annotated in UniProtKB. A match to an annotated site increases confidence in your identification. A novel site may represent a new biological finding, but it requires additional validation.

The precision of the text-mining extraction varies from 57 to 94 percent and recall from 75 to 95 percent, depending on the type of modification. This means that the PTM annotations in UniProtKB are not complete. Absence of a PTM annotation does not mean the modification does not occur. It may mean that the modification has not been reported in the literature or that the text-mining system did not capture the relevant publication.

### Expression Field

The Expression field describes where and when the protein is expressed. This information includes tissue-specific expression patterns, developmental stage expression, and induction conditions. For proteomics researchers, the Expression field helps contextualize protein identifications. A protein identified in a tissue where it is not expected to be expressed may indicate contamination, a novel expression pattern, or an annotation error.

The Expression field is more sparsely annotated than other fields because expression data is often scattered across many publications. The absence of expression information does not mean the protein is not expressed in your sample. It simply means that the information has not been curated into the database.

### Gene Ontology Annotations

Gene Ontology annotations are a structured vocabulary that describes three aspects of protein biology: molecular function, biological process, and cellular component. UniProtKB provides Gene Ontology annotations for most entries, and these annotations are widely used in downstream bioinformatics analysis.

The Gene Ontology annotations include evidence codes that indicate the basis for each annotation. Experimental evidence codes are stronger than computational evidence codes. When you perform enrichment analysis on your proteomics data, the quality of the Gene Ontology annotations directly affects the reliability of your results.

## Exporting Annotations for Downstream Analysis

The export step converts the annotation data from the UniProtKB website into a format that can be used in your analysis pipeline. The choice of export format depends on your downstream tools and the complexity of the analysis.

### Tab-Separated Values Export

The tab-separated values export is the most common format for proteomics researchers. This format produces a flat table with one row per protein and one column per selected annotation field. The table can be opened in spreadsheet software or imported into statistical analysis tools.

When you select columns for the tab-separated export, consider the structure of the data. Some fields contain multiple values separated by semicolons. For example, a protein may have multiple subcellular locations or multiple Gene Ontology terms. The semicolon-separated values require parsing before you can use them in enrichment analysis or other tools that expect one value per cell.

### FASTA Export

The FASTA export produces sequence records in the standard FASTA format. This format is useful when you need the protein sequences for downstream analysis such as motif searching or structural prediction. The FASTA header includes the accession and entry name, which allows you to map the sequences back to the annotation data.

### Programmatic Access Through the REST API

For large-scale or repeated analyses, the REST API provides programmatic access to UniProtKB data. The API allows you to retrieve annotations for thousands of accessions in a single request, and it supports the same column selection as the website interface. The API returns data in JSON or XML format, which can be parsed by scripting languages.

The REST API is documented on the UniProt website, and the documentation includes examples for common use cases. The SPARQL endpoint provides an additional programmatic access method that supports complex queries across the UniProtKB data model. The SPARQL endpoint is useful when you need to combine annotation data with other biological data sources.

## Practical Workflow for Proteomics Annotation Retrieval

The following workflow integrates the retrieval and interpretation steps into a reproducible process that can be applied to any proteomics experiment.

### Step 1: Generate the Accession List

Export the protein accession list from your proteomics search engine. The format depends on your search engine, but most tools provide a protein groups table that includes the leading protein accession for each group. Include the accession column in your export. If your search engine reports multiple accessions per protein group, decide whether to use the leading accession or all accessions. Using the leading accession simplifies the annotation retrieval but may miss isoform-specific annotations.

### Step 2: Map Accessions to UniProtKB

Submit the accession list to the Retrieve/ID Mapping tool. Select the input identifier type that matches your search engine output. Review the mapping results to identify accessions that failed to map. Failed mappings can occur for several reasons: the accession may be obsolete, the accession may belong to a different database, or the accession may contain a formatting error.

For failed mappings, check whether the accession exists in UniParc. If the accession is in UniParc but not in UniProtKB, the sequence has been excluded from the current UniProtKB release. This situation can occur when a sequence is no longer considered part of a reference proteome. You may need to use the UniParc record or find an alternative accession for the same protein.

### Step 3: Select Annotation Columns

Choose the annotation columns that match your analysis needs. For a standard proteomics analysis, include the following columns: Entry, Entry name, Reviewed, Protein names, Gene names, Organism, Length, Function, Subcellular location, PTM, Expression, and Gene ontology. Add columns for Cross-reference and Protein existence if you need those data.

The Reviewed column indicates whether the entry is in Swiss-Prot or TrEMBL. The Protein existence column indicates the type of evidence supporting the protein's existence. These columns help you assess the confidence level of the annotations.

### Step 4: Export and Validate the Data

Export the annotation table in tab-separated values format. Open the file in a spreadsheet or text editor and verify that the number of rows matches the number of successfully mapped accessions. Check that the accession column contains the expected values and that the annotation columns contain data instead of empty fields.

Validate the export by spot-checking several accessions against the UniProtKB website. Select a few proteins that are well characterized in your organism of interest and confirm that the exported annotations match the website entries. This validation step catches formatting errors or column misalignments before you proceed to downstream analysis.

### Step 5: Integrate Annotations into Your Analysis

Import the annotation table into your analysis pipeline. The integration approach depends on your tools. If you use R for statistical analysis, the Bioconductor project provides packages for working with protein annotation data. The Bioconductor project offers official package documentation, workflow guidance, and installation instructions for reproducible genomic analysis. If you use Python, the pandas library can read the tab-separated values file directly.

For enrichment analysis, convert the Gene Ontology annotations into the format expected by your enrichment tool. Most enrichment tools accept a list of Gene Ontology terms per protein or a binary matrix of protein-term associations. The conversion requires parsing the semicolon-separated Gene Ontology terms in the export file.

## Options and Tradeoffs in Annotation Retrieval

Different retrieval approaches have different strengths and weaknesses. The choice depends on the scale of your analysis, the reproducibility requirements, and your technical skills.

### Website Interface Versus Programmatic Access

The website interface is appropriate for small lists and exploratory analysis. It provides a visual interface that helps you understand the structure of the data. The website is also useful for spot-checking individual proteins and for learning the annotation fields.

Programmatic access through the REST API is appropriate for large lists and repeated analyses. The API supports automation, which improves reproducibility and reduces the risk of manual errors. The API also allows you to integrate annotation retrieval into your analysis pipeline, so the retrieval step is documented and versioned.

The tradeoff is the learning curve. The website interface requires no programming skills, while the REST API requires familiarity with scripting and HTTP requests. For researchers who are new to programming, the Carpentries lessons provide foundational training in computing and data skills that are useful for working with bioinformatics resources.

### Reviewed Versus Unreviewed Entries

The choice between reviewed and unreviewed entries affects the confidence level of your annotations. Reviewed entries have human-curated annotations that have been checked against the literature. Unreviewed entries have automated annotations that may be incomplete or incorrect.

For proteomics research, the practical approach is to use all entries but to track the review status. When you interpret your results, give more weight to annotations from reviewed entries. When you report your findings, note the review status of the proteins that support your conclusions.

The UniProtKB production pipeline is designed to limit the sequences available in UniProtKB to high-quality, non-redundant reference proteomes. This design choice means that the database increasingly contains well-supported sequences, which improves the overall quality of the annotation data.

### Manual Curation Versus Automated Annotation

The manual curation process ensures that Swiss-Prot entries contain the latest functional data. The UniProt consortium continues to manually curate the scientific literature to add new functional information. The consortium also encourages community curation to ensure that key publications are not missed.

The automated annotation methods used for TrEMBL entries predict information for unstudied proteins. These predictions are useful for proteins that lack experimental data, but they carry lower confidence than manually curated annotations. The automated methods use machine learning techniques to predict functional information, and the predictions are updated as new data becomes available.

## Records and Measurements for Annotation Retrieval

Keeping records of your annotation retrieval process supports reproducibility and troubleshooting. The following records should be maintained for each proteomics experiment.

### Retrieval Log

Document the date of retrieval, the UniProtKB release version, and the retrieval method. The release version matters because annotations change between releases. A protein that is unreviewed in one release may become reviewed in a later release after manual curation. A protein that is annotated with a phosphorylation site in one release may have additional sites added in a later release.

The retrieval log should include the input accession list, the mapping results, and the export file. Store these files in your project directory with clear naming conventions that include the retrieval date.

### Mapping Statistics

Record the number of input accessions, the number of successfully mapped accessions, and the number of failed mappings. Calculate the mapping rate as the percentage of input accessions that mapped to UniProtKB entries. A low mapping rate may indicate a problem with your accession format or a mismatch between your search engine database and the current UniProtKB release.

For failed mappings, record the accession and the reason for failure if you can determine it. Common reasons include obsolete accessions, accessions from a different database, and formatting errors. Track the failure reasons across experiments to identify systematic issues.

### Annotation Coverage

Record the percentage of mapped proteins that have annotations in each field of interest. For example, you might record the percentage of proteins with a Function annotation, the percentage with a Subcellular location annotation, and the percentage with a PTM annotation. The coverage varies by field and by organism. Some fields are more completely annotated than others.

The coverage measurement helps you interpret the results of downstream analysis. If only 50 percent of your proteins have PTM annotations, then a PTM enrichment analysis is based on a subset of your data. The coverage measurement tells you how much of your data is represented in each analysis.

## Common Failure Patterns and Troubleshooting

Several failure patterns recur when researchers retrieve annotations from UniProtKB. Recognizing these patterns helps you troubleshoot quickly and avoid wasted effort.

### Accession Format Errors

The most common failure is submitting accessions in the wrong format. Proteomics search engines sometimes append version numbers or species prefixes to accessions. The Retrieve/ID Mapping tool expects the accession without these additions. Strip version suffixes and species prefixes before submission.

Another format error is including header rows or descriptive text in the accession list. The mapping tool expects a plain list with one accession per line. Remove any headers, footers, or explanatory text before submission.

### Obsolete Accessions

Accessions can become obsolete when the underlying sequence is merged with another sequence or when the entry is deleted. The UniParc database stores obsolete data that has been excluded from UniProtKB. If an accession fails to map, check UniParc to determine whether the accession is obsolete and whether a replacement accession exists.

The UniProtKB release notes document accession changes. If you maintain a local database of accessions from previous experiments, periodically update the accessions to the current release to avoid obsolete accession errors.

### Taxonomy Mismatches

A taxonomy mismatch occurs when the accession list contains proteins from a different organism than expected. This situation can arise when a search engine database contains sequences from multiple organisms and the taxonomy filter was not applied correctly. The mapping tool will return entries for the correct organism, but the annotations may not match your experimental context.

Check the Organism column in the export file to verify that all mapped proteins belong to the expected species. If you find proteins from unexpected species, review your search engine settings and the original identification results.

### Column Selection Errors

The column selection interface can produce unexpected output if the selected columns do not match the analysis needs. For example, selecting the wrong Gene Ontology column may produce a file that cannot be parsed by your enrichment tool. Review the column names in the export file and verify that they match the expected format.

The UniProtKB website provides documentation for each column, including the format of the values. Consult the documentation when you are unsure about the structure of a column.

## Limitations of UniProtKB Annotations

UniProtKB annotations have limitations that affect their use in proteomics research. Understanding these limitations prevents overinterpretation of the data.

### Incomplete Annotation Coverage

Not all proteins have annotations in all fields. The annotation coverage varies by organism and by field. Well-studied organisms such as human and mouse have more complete annotations than less-studied organisms. Proteins that have been studied experimentally have more annotations than proteins that have only been predicted.

The absence of an annotation does not mean the protein lacks the corresponding feature. A protein without a PTM annotation may still have post-translational modifications that have not been reported in the literature. A protein without an Expression annotation may still be expressed in your sample.

### Annotation Errors

Annotations can be incorrect. Manual curation can introduce errors when the curator misinterprets the literature. Automated annotation can introduce errors when the prediction method is applied outside its valid domain. The evidence codes provide some indication of the confidence level, but they do not guarantee correctness.

When you use annotations to support a biological conclusion, verify the critical annotations against the primary literature. The literature citations in the entry provide the source for each annotation. Check the cited publications to confirm that the annotation matches the reported findings.

### Temporal Changes

Annotations change over time as new data becomes available. A protein that is annotated as having a single subcellular location may be re-annotated with multiple locations after new experiments are published. A protein that is unreviewed may become reviewed after manual curation.

The temporal nature of annotations means that your analysis results depend on the release version you used. If you repeat an analysis with a different release version, the results may differ. Document the release version in your records to support reproducibility.

## Quality Controls for Annotation-Based Analysis

Quality controls help you assess the reliability of your annotation-based analysis and identify potential problems before they affect your conclusions.

### Review Status Distribution

Calculate the percentage of your mapped proteins that are in Swiss-Prot versus TrEMBL. A high percentage of TrEMBL entries indicates that many of your identified proteins have not been manually curated. This situation is common for less-studied organisms and for novel proteins.

When you interpret your results, consider the review status distribution. Conclusions based primarily on TrEMBL annotations carry lower confidence than conclusions based on Swiss-Prot annotations. Report the review status distribution in your methods so that readers can assess the confidence level.

### Evidence Code Distribution

Examine the evidence codes for the annotations you use in your analysis. Experimental evidence codes indicate that the annotation is supported by direct experimental data. Computational evidence codes indicate that the annotation is predicted. The distribution of evidence codes tells you how much of your analysis is based on experimental data versus predictions.

For critical conclusions, require experimental evidence. For exploratory analysis, computational evidence may be acceptable, but the limitations should be acknowledged.

### Cross-Validation with Other Databases

Cross-validate your annotations against other databases to identify discrepancies. The NCBI data resources provide complementary protein information that can be used for cross-validation. If a protein has conflicting annotations in UniProtKB and NCBI, investigate the source of the conflict before drawing conclusions.

The cross-validation step is particularly important for proteins that are central to your conclusions. A protein that is annotated as a kinase in UniProtKB but as a phosphatase in NCBI requires careful investigation to resolve the discrepancy.

## Professional Escalation Criteria

Some situations require escalation beyond the standard annotation retrieval workflow. Recognize these situations and seek additional expertise when they arise.

### Persistent Mapping Failures

If a significant percentage of your accessions fail to map and you cannot resolve the failures through UniParc or format corrections, escalate the issue. The problem may indicate a systematic issue with your search engine database or a fundamental mismatch between your identification pipeline and the current UniProtKB release.

Consult the UniProtKB documentation and the search engine documentation to identify the source of the mismatch. If the issue persists, contact the search engine vendor or the UniProt help desk for assistance.

### Conflicting Annotations for Critical Proteins

If you identify conflicting annotations for a protein that is central to your conclusions, escalate the issue to a domain expert. The conflict may indicate an annotation error, a genuine biological complexity, or a misinterpretation on your part. A domain expert can review the primary literature and provide an informed assessment.

The UniProtKB entry includes literature citations that provide the source for each annotation. Review the cited publications to understand the evidence for each conflicting annotation.

### Regulatory or Clinical Implications

If your proteomics analysis has regulatory or clinical implications, escalate the annotation interpretation to the appropriate authority. Protein annotations that support diagnostic, prognostic, or therapeutic decisions require careful validation beyond what UniProtKB provides.

The UniProtKB database has been awarded Global Core Biodata Resource status in recognition of its value to the scientific community. This status reflects the importance of the database, but it does not replace the need for domain-specific expertise in clinical or regulatory contexts.

## Training and Skill Development for Annotation Retrieval

Developing proficiency in annotation retrieval requires training in both the specific tools and the underlying bioinformatics concepts. Several training resources provide structured learning pathways.

### EMBL-EBI Training Resources

The European Bioinformatics Institute provides training resources for bioinformatics data resources and practical analysis education. The EMBL-EBI training materials cover the use of biological databases, including UniProtKB, and provide hands-on exercises that build practical skills. The training resources are appropriate for researchers at all levels, from beginners to advanced users.

### Galaxy Training Network

The Galaxy Training Network provides accessible workflow training and analysis tutorials. The Galaxy platform supports reproducible analysis through a web-based interface, and the training materials cover a range of bioinformatics topics. The Galaxy Training Network tutorials include practical exercises that teach annotation retrieval and downstream analysis within the Galaxy environment.

### nf-core Documentation

The nf-core community provides pipeline standards, usage documentation, and configuration guidance for reproducible workflows. The nf-core documentation is relevant for researchers who use Nextflow pipelines for proteomics analysis. The documentation covers the standards for pipeline development and usage, which helps researchers integrate annotation retrieval into their existing workflows.

### Bioconductor Documentation

The Bioconductor project provides official package documentation, workflow guidance, and installation instructions for reproducible genomic analysis. The Bioconductor packages for proteomics analysis include tools for working with protein annotation data. The documentation includes examples that demonstrate how to retrieve and interpret annotations within the R environment.

### The Carpentries Lessons

The Carpentries lessons provide foundational training in computing and data skills. The lessons cover shell, Git, and programming concepts that are useful for bioinformatics work. The Carpentries lessons are appropriate for researchers who need to build the technical skills required for programmatic annotation retrieval.

## Frequently Asked Questions

### What is the difference between UniProtKB/Swiss-Prot and UniProtKB/TrEMBL?

UniProtKB/Swiss-Prot contains manually annotated entries that have been curated by biologists who read the scientific literature and add functional information. UniProtKB/TrEMBL contains computer-annotated entries that are produced through automated prediction methods. The review status of an entry indicates which section it belongs to, and this status affects the confidence level of the annotations.

### How do I retrieve annotations for a large list of protein accessions?

Use the Retrieve/ID Mapping tool on the UniProtKB website for lists of up to several hundred accessions. For larger lists or repeated analyses, use the REST API, which supports programmatic retrieval of annotations for thousands of accessions in a single request. The REST API returns data in JSON or XML format that can be parsed by scripting languages.

### What should I do when an accession fails to map in UniProtKB?

Check whether the accession exists in UniParc, which stores all publicly available protein sequence data including obsolete data excluded from UniProtKB. If the accession is in UniParc, the sequence has been excluded from the current UniProtKB release. Verify the accession format and check for version suffixes or species prefixes that may need to be removed.

### How do I know whether an annotation is based on experimental evidence or prediction?

UniProtKB provides evidence codes for annotations that indicate the type of evidence supporting each claim. Experimental evidence codes indicate that the annotation is supported by direct experimental data. Computational evidence codes indicate that the annotation is predicted by automated methods. The evidence codes are displayed on the entry page and included in the export file.

### Can I use UniProtKB annotations for enrichment analysis?

Yes, the Gene Ontology annotations in UniProtKB are commonly used for enrichment analysis. Export the Gene Ontology annotations in tab-separated values format and convert them to the format expected by your enrichment tool. The conversion requires parsing the semicolon-separated Gene Ontology terms in the export file.

### How often is UniProtKB updated?

UniProtKB is updated regularly, with new releases published every two weeks. The release notes document the changes in each release, including new entries, updated annotations, and accession changes. Document the release version in your records to support reproducibility of your analysis.

### What is the Rhea knowledgebase and how does it relate to UniProtKB?

Rhea is a comprehensive expert-curated knowledgebase of biochemical reactions that describes reaction participants using the ChEBI ontology. UniProtKB uses Rhea as the standard for annotation of enzymatic reactions. The Rhea integration provides improved search and query facilities for the UniProt website, REST API, and SPARQL endpoint.

### How do I cite UniProtKB in my publications?

The UniProtKB publication in Nucleic Acids Research describes the database and its ongoing updates. The UniProt consortium also maintains a website that provides citation guidance. When you use UniProtKB data in your publications, cite the database and note the release version you used for your analysis.

## Related Bioinformatics Guides

- [Proteomics Analysis Tools: A Comparative Guide for Functional Interpretation](/knowledge/bioinformatics/proteomics-analysis-tools-a-comparative-guide-for-functional-interpretation)
- [Functional Annotation of Metagenomes: A Guide to Databases and Pipelines](/knowledge/bioinformatics/functional-annotation-of-metagenomes-a-guide-to-databases-and-pipelines)
- [Volcano Plot Proteomics: How to Create and Interpret Them Effectively](/knowledge/bioinformatics/volcano-plot-proteomics-how-to-create-and-interpret-them-effectively)
- [AlphaFold2-Based Structural Modeling and Functional Annotation of PRRSV Nonstructural Proteins](/knowledge/bioinformatics/alphafold2-structural-modeling-prrsv-nonstructural-proteins)
- [How to Interpret Gene Set Enrichment Analysis Results](/knowledge/bioinformatics/how-to-interpret-gene-set-enrichment-analysis-results)

## References and Further Reading

- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Bioconductor](https://bioconductor.org/). Bioconductor Project.
- [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.
- [nf-core Documentation](https://nf-co.re/docs). nf-core.
- [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.
- [UniProt: the Universal Protein Knowledgebase in 2025.](https://pubmed.ncbi.nlm.nih.gov/39552041). Nucleic acids research, 2025.
- [UniProtKB/Swiss-Prot.](https://pubmed.ncbi.nlm.nih.gov/18287689). Methods in molecular biology (Clifton, N.J.), 2007.
- [Enzyme annotation in UniProtKB using Rhea.](https://pubmed.ncbi.nlm.nih.gov/31688925). Bioinformatics (Oxford, England), 2020.
- [Searching and Navigating UniProt Databases.](https://pubmed.ncbi.nlm.nih.gov/36912607). Current protocols, 2023.
- [Application of text-mining for updating protein post-translational modification annotation in UniProtKB.](https://pubmed.ncbi.nlm.nih.gov/23517090). BMC bioinformatics, 2013.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.