# How to Submit Proteomics Data to PRIDE: A Step-by-Step Guide to Using the ProteomeXchange Submission Pipeline

Submitting mass spectrometry proteomics data to the PRIDE database at the European Bioinformatics Institute requires preparation of raw instrument files, peak list files, search result files, and complete metadata annotation through the ProteomeXchange submission pipeline. This guide provides the practical steps for depositing MS-based proteomics data, covering the two available workflows, file requirements, metadata templates, and common pitfalls that delay or reject submissions. The target reader is a biology student, researcher, or laboratory professional who has generated MS-based proteomics data and needs to deposit it for publication, funding compliance, or community reuse.

## Understanding PRIDE and the ProteomeXchange Consortium

PRIDE, which stands for PRoteomics IDEntifications database, is the leading mass spectrometry based proteomics data repository and a founding member of the ProteomeXchange consortium. The ProteomeXchange consortium was established to standardize and facilitate the submission and dissemination of MS-based proteomics data in the public domain. Within this consortium, PRIDE acts as the initial submission point for MS/MS datasets. The consortium includes six member resources: PRIDE, PeptideAtlas including PASSEL, MassIVE, jPOST, iProX, and Panorama Public. Each resource accepts submissions, but the workflow and file requirements differ slightly between them. This guide focuses on the PRIDE submission path because PRIDE is the most commonly used repository for European researchers and serves as the default submission point for many journals.

The number of datasets submitted to PRIDE Archive has grown substantially. As of the 2025 update, PRIDE receives on average around 534 datasets per month. This volume has been made possible by continuous infrastructure improvements, including a new file transfer protocol for very large datasets called Globus, a new data resubmission pipeline, and an automatic dataset validation process. Understanding these systems before you start will reduce the time spent troubleshooting during submission.

The ProteomeXchange consortium has been operating for over a decade, and as of June 2022, more than 34,233 datasets had been submitted to PX resources, with 20,062 of those submitted in the preceding three years. This growth demonstrates the increasing expectation from journals and funding bodies that proteomics data be made publicly available. The consortium has also developed Universal Spectrum Identifiers and improved the capture of experimental metadata annotations to support data reuse activities.

## Preparing Your Data Files Before Submission

The first and most important step in the submission process is organizing your files. PRIDE requires three categories of files for a complete submission: raw instrument files, peak list files, and result files. Each category serves a distinct purpose in the data validation and reanalysis process.

### Raw Instrument Files

Raw files are the unprocessed output from your mass spectrometer. These files contain the full record of all detected ions, their intensities, and their fragmentation patterns. The format depends on your instrument vendor. Thermo Fisher instruments produce .raw files, Bruker instruments produce .d folders, and SCIEX instruments produce .wiff files. PRIDE accepts all vendor formats, but you must ensure that the files are not corrupted and that you have the original unprocessed versions. Do not submit files that have been centroided or otherwise processed unless you also include the original raw data.

The size of raw files is often the main practical challenge. A single LC-MS/MS run can produce files ranging from hundreds of megabytes to several gigabytes depending on the instrument, the gradient length, and the acquisition mode. For a typical project with dozens of runs, the total data volume can reach hundreds of gigabytes or even terabytes. Before you begin the submission process, calculate the total size of your raw files and verify that you have sufficient storage and bandwidth to transfer them. PRIDE has implemented the Globus file transfer protocol specifically to handle very large datasets, so you should plan to use this method if your total data volume exceeds what a standard browser upload can handle.

### Peak List Files

Peak list files contain the processed spectra that were used for database searching. Common formats include MGF, mzXML, mzML, and MS2. These files are derived from the raw data by software that extracts the precursor masses and fragment ion information. The peak list files are essential because they allow reviewers and reanalysts to verify the identification results without needing to reprocess the raw files.

When preparing peak list files, verify that they were generated with the same parameters that you used for your database search. If you used a search engine such as Mascot, Sequest, or MaxQuant, the peak list files should match the search input exactly. Some search engines generate their own peak lists internally, in which case you need to export them in a standard format. Check that the peak list files are complete and that no spectra were lost during conversion.

### Result Files

Result files contain the peptide and protein identifications from your database search. The format depends on the search engine and the downstream analysis software. Common result file formats include .dat files from Mascot, .msf files from Proteome Discoverer, .xlsx or .txt files from MaxQuant, and .pep.xml or .prot.xml files from Trans-Proteomic Pipeline. PRIDE accepts these formats and converts them into a standard representation during the submission process.

For quantitative proteomics experiments, you must also include the quantification information. This may be embedded in the result files or provided as separate tables. If you used label-free quantification, include the intensity values for each peptide and protein across all samples. If you used isobaric labeling such as TMT or iTRAQ, include the reporter ion intensities. The quantification data must be complete and consistent with the identification data.

## Choosing Between Complete and Partial Submission Workflows

The ProteomeXchange consortium offers two submission workflows: complete and partial. The choice between them depends on the nature of your data and the requirements of your target journal.

### Complete Submission Workflow

A complete submission includes all three file categories: raw files, peak list files, and result files. This workflow is required for most journals that mandate data deposition. Complete submissions receive a ProteomeXchange accession number that can be cited in your manuscript. The data are fully available to the community, and other researchers can reanalyze your raw data with different search parameters or software.

The complete workflow is appropriate when your data are not subject to confidentiality restrictions and when you have the storage and bandwidth to transfer all files. Most standard proteomics projects, including discovery proteomics, quantitative comparisons, and post-translational modification analyses, should use the complete workflow.

### Partial Submission Workflow

A partial submission includes only the peak list files and result files, without the raw instrument files. This workflow is appropriate when the raw data cannot be shared due to patient confidentiality, commercial restrictions, or other constraints. Partial submissions receive a different type of accession and the data availability is more limited.

The partial workflow is also useful for very large datasets where transferring the raw files is impractical. However, you should be aware that many journals require complete submissions, and reviewers may request raw data during the review process. If you choose the partial workflow, document the reason clearly in your submission and be prepared to justify it to editors and reviewers.

## Using PRIDE Converter for File Conversion and Metadata Entry

PRIDE Converter is a graphical user interface tool that simplifies the submission process by converting your result files into the PRIDE XML format and guiding you through metadata entry. The tool was developed to address the growing need for public data availability and the increasing journal requirements for MS data deposition. While the original PRIDE Converter has been superseded by newer tools, the principles it established remain central to the submission process.

### Installing and Running PRIDE Converter

PRIDE Converter runs on any system with Java installed. Download the tool from the official PRIDE website and launch it on your local machine. The tool presents a step by step wizard that guides you through the process of selecting your result files, mapping the search engine output to the PRIDE XML schema, and entering the experimental metadata.

The tool accepts result files from the major search engines including Mascot, Sequest, and others. When you load your result files, the tool parses the identifications and displays them in a table. You can review the identifications and remove any that are decoy hits or contaminants before conversion. This step is important because it ensures that the submitted data are clean and that the false discovery rate is properly controlled.

### Mapping Search Engine Output to PRIDE XML

The conversion process requires mapping the fields in your search engine output to the PRIDE XML schema. This mapping is mostly automatic, but you should verify that the key fields are correctly assigned. The essential fields include the peptide sequence, the precursor charge and mass, the modifications, the protein accession, and the scores. If any fields are missing or incorrectly mapped, the conversion will fail or produce incomplete data.

For quantitative data, PRIDE Converter supports the inclusion of quantification values. The tool can parse quantification information from the search engine output or from separate files. Verify that the quantification values are correctly associated with the correct peptides and samples.

### Entering Experimental Metadata

The metadata entry step is where many submissions encounter problems. PRIDE requires detailed metadata about the experiment, including the species, tissue, cell line, disease state, instrument, fragmentation method, search engine, and database used. This information is stored in the PRIDE XML file and is used by the repository to index and display your dataset.

The metadata must be entered in a structured format. PRIDE uses the MAGE-TAB format for metadata annotation, which is a tab-delimited format originally developed for microarray data but adapted for proteomics. The MAGE-TAB format includes a sample description file, a data file description, and an experiment description. Each sample must be described with attributes such as organism, tissue, cell type, and any relevant experimental conditions.

The importance of complete metadata cannot be overstated. A review of human tissue datasets submitted to ProteomeXchange repositories found that about half of the retrieved datasets lacked the annotations or metadata necessary to directly replicate the analysis. This lack of metadata poses a serious challenge to data reusability. The review highlighted the need to increase awareness of the MAGE-TAB file format for metadata in the community. When you submit your data, treat the metadata as equally important as the data files themselves.

## Step-by-Step Submission Process

The submission process follows a defined sequence of steps. Following this sequence in order will minimize errors and reduce the time required for your dataset to become public.

### Step 1: Create an Account and Log In

Before you can submit data to PRIDE, you need an account. Go to the PRIDE website and register with your institutional email address. The registration process is straightforward and requires basic information about you and your affiliation. Once your account is created, log in to the submission portal.

### Step 2: Start a New Submission

From the submission portal, select the option to start a new submission. You will be asked to choose between the complete and partial workflows. Select the workflow that matches your data availability and journal requirements. You will also be asked to provide a title for your dataset and a brief description of the experiment.

### Step 3: Upload Your Files

The file upload step is where you transfer your raw files, peak list files, and result files to the PRIDE server. For small datasets, you can use the browser based upload. For larger datasets, use the Globus file transfer protocol, which is designed for very large files and provides faster and more reliable transfer than browser uploads.

During the upload, verify that all files are transferred completely. The submission portal shows the status of each file. If any file fails to upload, check the file integrity and retry. Do not proceed to the next step until all files are uploaded successfully.

### Step 4: Enter Metadata

After the files are uploaded, you will be guided through the metadata entry process. This includes the experiment description, sample annotations, and protocol information. The metadata entry forms vary depending on the workflow you selected. For complete submissions, you must provide detailed information about the instrument settings, the search parameters, and the database used.

Use the MAGE-TAB format for sample annotations. Create a table where each row represents a sample and each column represents an attribute. Include the organism, tissue, cell line, disease state, and any treatment conditions. The more complete your sample annotations, the more likely your data will be reusable by other researchers.

### Step 5: Validate and Submit

Once the files are uploaded and the metadata is entered, the submission system runs an automatic validation process. This process checks that the files are in the correct format, that the metadata is complete, and that the data are internally consistent. The validation process was implemented as part of the infrastructure improvements to PRIDE and helps to reduce the number of submissions that require manual intervention.

If the validation fails, the system will report the errors. Common errors include missing metadata fields, incorrectly formatted files, and inconsistencies between the peak list files and the result files. Fix the reported errors and resubmit. If the validation passes, you can submit your dataset for curation.

### Step 6: Curation and Accession

After submission, your dataset enters the curation queue. A PRIDE curator reviews your submission to verify that the data are complete and that the metadata is accurate. The curator may contact you with questions or requests for additional information. Respond promptly to these requests to avoid delays.

Once the curation is complete, your dataset receives a ProteomeXchange accession number in the format PXD followed by six digits. This accession number is what you cite in your manuscript. The dataset becomes public immediately or after a specified embargo period, depending on your preferences and the journal requirements.

## At a Glance: Submission Workflow Comparison

| Workflow | Files Required | Accession Type | Data Availability | Best For |
|----------|---------------|----------------|-------------------|----------|
| Complete | Raw files, peak lists, result files | PXD accession | Full public access | Standard discovery and quantitative proteomics |
| Partial | Peak lists, result files | PXD accession with restricted raw data | Limited access | Confidential or sensitive data |
| Resubmission | Updated files and metadata | Original PXD accession | Updated public access | Correcting or extending published datasets |

## Metadata Annotation Standards and Best Practices

The metadata you provide with your dataset determines whether other researchers can understand and reuse your data. The ProteomeXchange consortium has made improvements in capturing experimental metadata annotations, but the responsibility for providing complete metadata rests with the submitting researcher.

### Required Metadata Fields

PRIDE requires certain metadata fields for every submission. These include the experiment title, a description of the experimental design, the species studied, the tissue or cell type, the instrument used, the fragmentation method, the search engine, and the sequence database. For quantitative experiments, you must also describe the quantification method and the normalization approach.

The experiment description should be written in clear language that a researcher outside your immediate field can understand. Describe the biological question, the experimental design, the number of replicates, and the key findings. Avoid jargon and undefined abbreviations.

### Sample Annotations

Each sample in your dataset must be annotated with its biological attributes. Use controlled vocabulary terms where available. For example, use the NCBI taxonomy for species names and the Cell Line Ontology for cell lines. The use of controlled vocabulary ensures that your samples can be compared with samples from other datasets.

The MAGE-TAB format provides a structured way to record sample annotations. Create a sample description file where each row is a sample and each column is an attribute. Include the biological replicate information, the technical replicate information, and any treatment or condition applied. The more detailed your sample annotations, the more valuable your dataset becomes for reanalysis.

### Protocol Information

Describe the experimental protocols in sufficient detail that another researcher could reproduce your experiment. This includes the sample preparation protocol, the chromatography conditions, the mass spectrometry acquisition parameters, and the data analysis parameters. The protocol information can be provided as free text or as structured fields, depending on the submission interface.

For the mass spectrometry parameters, include the instrument model, the ionization source, the mass analyzer, the fragmentation method, the resolution settings, and the scan range. For the data analysis parameters, include the search engine, the search engine version, the precursor mass tolerance, the fragment mass tolerance, the enzyme specificity, the number of missed cleavages, the fixed and variable modifications, and the target false discovery rate.

## File Formats and Conversion Tools

The success of your submission depends on using the correct file formats. PRIDE accepts a range of formats for each file category, and the submission system validates the format during the upload process.

### Accepted Raw File Formats

PRIDE accepts raw files in the native vendor formats. This includes Thermo .raw, Bruker .d, SCIEX .wiff, Agilent .d, and Waters .raw. The repository does not require conversion of raw files to a standard format because the vendor formats contain the complete instrument data and are the most faithful representation of the original acquisition.

If your instrument produces files in a format that PRIDE does not accept, you may need to convert the files to mzML or another open format. The conversion should preserve all the information in the original files, including the ion mobility data if applicable. Verify the converted files by comparing the number of spectra and the total ion current with the original files.

### Accepted Peak List Formats

The peak list files must be in a standard format that the search engines and the repository can read. The most common formats are MGF, mzXML, mzML, and MS2. MGF is the most widely supported format and is produced by most peak list extraction tools. mzML is the HUPO PSI standard format and is recommended for new submissions.

When converting peak lists to a standard format, verify that the precursor masses, charge states, and fragment ion masses are preserved. Check that the retention times are included if your downstream analysis requires them. The peak list files should contain all the spectra that were used for the database search, without any filtering or removal.

### Accepted Result File Formats

The result files must be in a format that PRIDE can parse and convert to its internal representation. The accepted formats include Mascot .dat, Proteome Discoverer .msf, MaxQuant evidence and protein groups tables, and Trans-Proteomic Pipeline .pep.xml and .prot.xml files. If your search engine produces a format that is not directly accepted, you may need to convert the results to one of the accepted formats.

For quantitative data, the result files must include the quantification values. The format of the quantification data depends on the quantification method. For label-free quantification, include the intensity values for each peptide and protein. For isobaric labeling, include the reporter ion intensities. For SILAC, include the heavy to light ratios.

## Common Failure Patterns and How to Avoid Them

Many submissions encounter problems that delay publication or require resubmission. Understanding the common failure patterns will help you avoid these issues.

### Incomplete Metadata

The most common reason for submission delays is incomplete metadata. Reviewers and curators need to understand your experiment to evaluate your data. If the metadata is incomplete, the curator will request additional information, which adds time to the submission process.

To avoid this problem, prepare your metadata before you start the submission. Create a document that contains all the required fields and fill it in as you plan your experiment. Use the MAGE-TAB format for sample annotations and include all the protocol details. The time spent preparing metadata is much less than the time spent responding to curator requests.

### File Format Errors

File format errors occur when the files do not match the expected format or are corrupted during transfer. The automatic validation process catches many of these errors, but some may not be detected until the curation stage.

To avoid file format errors, verify your files before upload. Open the peak list files and check that they contain the expected number of spectra. Open the result files and check that the identifications are present. Verify that the raw files are not corrupted by checking their file sizes and checksums.

### Inconsistent Data

Inconsistencies between the peak list files and the result files can cause validation failures. For example, if the result files contain identifications for spectra that are not present in the peak list files, the validation will fail. This can happen if the peak list files were generated with different parameters than the search.

To avoid inconsistencies, generate the peak list files and the result files in the same analysis run. Use the same parameters for peak list extraction and database searching. If you need to regenerate the peak list files, rerun the database search with the new peak lists.

### Oversized File Transfers

Large datasets can fail to upload if the transfer method is not appropriate for the file sizes. Browser uploads are suitable for datasets up to a few gigabytes, but larger datasets require the Globus transfer protocol. The Globus protocol is designed for very large files and provides faster and more reliable transfer.

To avoid transfer failures, estimate the total size of your dataset before you start. If the total size exceeds 10 gigabytes, use Globus. If you are unsure, contact the PRIDE helpdesk for guidance.

## Records and Measurements to Keep During Submission

Maintaining detailed records of your submission process is important for reproducibility and for responding to curator requests.

### Submission Log

Keep a log of your submission activities, including the date and time of each step, the files uploaded, and any errors encountered. This log will help you track the progress of your submission and identify any steps that need to be repeated.

### File Checksums

Record the checksums of all your files before upload. The checksums allow you to verify that the files were transferred without corruption. If the curator reports a problem with a file, you can compare the checksum of the uploaded file with the checksum of your local copy.

### Metadata Version Control

Keep a versioned record of your metadata. If you need to update the metadata after submission, you can track the changes and ensure that the final version is consistent. The PRIDE resubmission pipeline allows you to update datasets, but you should document the changes you make.

## Quality Controls and Validation Checks

The PRIDE submission system includes automatic validation checks that verify the integrity and consistency of your submission. Understanding these checks will help you prepare a submission that passes validation on the first attempt.

### File Integrity Checks

The validation process checks that all files are present and that they are in the correct format. The system verifies the file sizes and checksums to ensure that the files were not corrupted during transfer. If a file is missing or corrupted, the validation will fail and you will need to reupload the file.

### Metadata Completeness Checks

The validation process checks that all required metadata fields are filled in. The system verifies that the sample annotations are present and that the protocol information is complete. If any required field is missing, the validation will fail and you will need to add the missing information.

### Data Consistency Checks

The validation process checks that the data are internally consistent. The system verifies that the peak list files contain the spectra that are referenced in the result files. The system also checks that the quantification values are consistent with the identification data. If any inconsistency is found, the validation will fail and you will need to correct the data.

## Training Resources for Proteomics Data Submission

Several training resources are available to help researchers learn the skills needed for successful proteomics data submission and analysis. These resources cover topics from basic bioinformatics to advanced workflow development.

### EMBL-EBI Training

The European Bioinformatics Institute offers training courses and online materials covering data resources and practical analysis education. These training opportunities are valuable for researchers who want to understand the PRIDE submission process in the context of the broader bioinformatics landscape. The training materials cover data resource usage, submission best practices, and analysis workflows.

### Galaxy Training Network

The Galaxy Training Network provides accessible workflow training and analysis tutorials that emphasize reproducibility. These tutorials are useful for researchers who want to learn how to process proteomics data before submission. The network offers hands-on exercises that demonstrate how to use Galaxy for mass spectrometry data analysis, including peptide identification and quantification workflows.

### nf-core Documentation

The nf-core community provides documentation for reproducible bioinformatics pipelines. These pipelines follow community standards for configuration and usage, making them useful for researchers who want to process their proteomics data consistently before submission. The documentation covers pipeline installation, configuration, and execution.

### The Carpentries Lessons

The Carpentries offers foundational lessons in computing, data, shell, Git, and programming. These lessons are valuable for researchers who need to develop the computational skills required for managing and submitting large proteomics datasets. The lessons cover file management, data organization, and version control, which are all relevant to the submission process.

### Bioconductor Documentation

Bioconductor provides official documentation for packages, workflows, installation, and reproducible genomic analysis. For proteomics researchers, Bioconductor offers packages for mass spectrometry data processing and analysis. The documentation covers package installation and usage, which is useful for researchers who want to analyze their data before submission.

## Limitations and Interpretation Boundaries

Understanding the limitations of the PRIDE submission process and the data it stores will help you use the repository appropriately.

### Data Reusability Depends on Metadata Quality

The reusability of your data depends on the quality of your metadata. A review of human tissue datasets submitted to ProteomeXchange repositories found that about half of the datasets lacked the annotations necessary to directly replicate the analysis. This finding highlights the importance of providing complete metadata with your submission. If you do not provide sufficient metadata, other researchers will not be able to use your data, and the value of your deposition is diminished.

### Raw Data Interpretation Requires Expertise

The raw data files in PRIDE are complex and require specialized software and expertise to interpret. Other researchers who download your raw files will need to process them with their own analysis pipelines. The peak list files and result files provide a more accessible entry point for reanalysis, but they still require bioinformatics expertise to interpret.

### Repository Growth Affects Search and Retrieval

The number of datasets in PRIDE continues to grow, with an average of around 534 datasets submitted per month. This growth makes it increasingly important to provide accurate and complete metadata so that your dataset can be found and retrieved by other researchers. The search functionality in PRIDE relies on the metadata you provide, so the quality of your metadata directly affects the discoverability of your data.

### Human Tissue Data Representation

Human tissue data are variably represented in the ProteomeXchange repositories, ranging between 10% and 25% of the total data submitted. Cancers are the most represented condition, followed by neuronal and cardiovascular diseases. If you are submitting human tissue data, be aware that the metadata requirements are particularly important for enabling meaningful reanalysis and comparison with other datasets.

## Professional Escalation Criteria

Knowing when to escalate a submission problem to the PRIDE helpdesk or to a senior colleague can save time and reduce frustration.

### When to Contact the PRIDE Helpdesk

Contact the PRIDE helpdesk if you encounter a technical problem that you cannot resolve on your own. This includes file upload failures, validation errors that you cannot fix, and questions about the submission process. The helpdesk can provide guidance on file formats, metadata requirements, and transfer protocols.

### When to Consult a Senior Colleague

Consult a senior colleague or your institutional bioinformatics support if you are unsure about the appropriate workflow for your data. This includes questions about whether to use the complete or partial workflow, how to handle sensitive data, and how to respond to curator requests. A senior colleague with experience in proteomics data deposition can provide valuable guidance.

### When to Request a Resubmission

If you discover an error in your submitted dataset, you can request a resubmission. The PRIDE resubmission pipeline allows you to update your dataset with corrected files or metadata. This is appropriate when you find an error in the identifications, when you need to add data, or when you need to correct the metadata. Document the changes you make and inform the curator of the reason for the resubmission.

## Frequently Asked Questions

### What is the difference between a complete and a partial submission to PRIDE?

A complete submission includes raw instrument files, peak list files, and result files, and results in a ProteomeXchange accession that provides full public access to all data. A partial submission includes only peak list files and result files, without the raw instrument files. Partial submissions are appropriate when raw data cannot be shared due to confidentiality or commercial restrictions, but many journals require complete submissions.

### How long does the PRIDE submission process take?

The time required depends on the size of your dataset and the completeness of your metadata. File upload time depends on the total data volume and the transfer method. After upload, the automatic validation process runs and may require corrections. The curation process by a PRIDE curator can take several days to several weeks depending on the workload. Responding promptly to curator requests will reduce the total time.

### What file formats does PRIDE accept for raw data?

PRIDE accepts raw files in the native vendor formats, including Thermo .raw, Bruker .d, SCIEX .wiff, Agilent .d, and Waters .raw. If your instrument produces a format that is not accepted, you may need to convert the files to mzML or another open format. Verify the converted files to ensure that all information is preserved.

### Can I submit quantitative proteomics data to PRIDE?

Yes, PRIDE accepts quantitative proteomics data. You must include the quantification values in your result files or as separate tables. The quantification method must be described in the metadata, including the type of quantification, the normalization approach, and the software used. The quantification values must be consistent with the identification data.

### How do I cite my PRIDE dataset in a manuscript?

You cite your dataset using the ProteomeXchange accession number in the format PXD followed by six digits. Include the accession number in the data availability statement of your manuscript. The accession number allows readers to access your dataset directly from the ProteomeXchange website.

### What should I do if my submission fails validation?

If your submission fails validation, the system will report the errors. Review the errors and fix the reported issues. Common errors include missing metadata fields, incorrectly formatted files, and inconsistencies between the peak list files and the result files. After fixing the errors, resubmit the dataset. If you cannot resolve the errors, contact the PRIDE helpdesk for assistance.

### Can I update my dataset after it has been submitted?

Yes, PRIDE provides a resubmission pipeline that allows you to update your dataset. You can correct errors in the files or metadata, add new data, or extend the dataset. Document the changes you make and inform the curator of the reason for the resubmission. The updated dataset retains the original accession number.

### What metadata is required for a PRIDE submission?

PRIDE requires metadata about the experiment title, experimental design, species, tissue or cell type, instrument, fragmentation method, search engine, and sequence database. For quantitative experiments, you must also describe the quantification method and normalization approach. Sample annotations should use the MAGE-TAB format with controlled vocabulary terms where available.

## Related Bioinformatics Guides

- [Proteomics Mass Spectrometry: From Sample Preparation to Data Analysis](/knowledge/bioinformatics/proteomics-mass-spectrometry-from-sample-preparation-to-data-analysis)
- [Mass Spectrometry-Based Proteomics: Data Analysis Pipelines and Tools](/knowledge/bioinformatics/mass-spectrometry-based-proteomics-data-analysis-pipelines-and-tools)
- [Spatial Proteomics Mass Spectrometry: Techniques and Applications](/knowledge/bioinformatics/spatial-proteomics-mass-spectrometry-techniques-and-applications)
- [Genomic Data Analysis Tools: A Comparative Guide for Researchers](/knowledge/bioinformatics/genomic-data-analysis-tools-a-comparative-guide-for-researchers)
- [Olink Proteomics: A Practical Guide to Panel Selection and Data Interpretation](/knowledge/bioinformatics/olink-proteomics-a-practical-guide-to-panel-selection-and-data-interpretation)

## References and Further Reading

- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Bioconductor](https://bioconductor.org/). Bioconductor Project.
- [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.
- [nf-core Documentation](https://nf-co.re/docs). nf-core.
- [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.
- [The PRIDE database at 20 years: 2025 update.](https://pubmed.ncbi.nlm.nih.gov/39494541). Nucleic acids research, 2025.
- [How to submit MS proteomics data to ProteomeXchange via the PRIDE database.](https://pubmed.ncbi.nlm.nih.gov/25047258). Proteomics, 2014.
- [Tissue proteomics repositories for data reanalysis.](https://pubmed.ncbi.nlm.nih.gov/37534389). Mass spectrometry reviews, 2024.
- [The ProteomeXchange consortium at 10 years: 2023 update.](https://pubmed.ncbi.nlm.nih.gov/36370099). Nucleic acids research, 2023.
- [Submitting proteomics data to PRIDE using PRIDE Converter.](https://pubmed.ncbi.nlm.nih.gov/21082439). Methods in molecular biology (Clifton, N.J.), 2011.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.