# STRING vs. BioGRID: Choosing the Right Protein-Protein Interaction Database for Your Proteomics Analysis

Proteomics experiments generate long lists of differentially expressed proteins, but those lists only become biologically meaningful when you map them onto protein-protein interaction (PPI) networks. STRING and BioGRID are two of the most commonly used databases for this purpose, yet they answer different questions and draw on different evidence types. STRING aggregates predicted and experimentally derived interactions across many organisms and assigns confidence scores, while BioGRID curates experimentally verified interactions from the primary literature with a focus on model organisms. Your choice between them depends on whether you need broad functional context for a discovery-oriented proteomics screen or high-confidence physical interaction data for a targeted mechanistic study. This article compares the two databases across data content, confidence scoring, coverage, and usability, and provides concrete decision criteria based on your research goals.

## Understanding the Core Differences Between STRING and BioGRID

The fundamental distinction between STRING and BioGRID lies in their evidence philosophy. STRING is a meta-database that integrates multiple sources of evidence, including experimental data, curated databases, text mining of the literature, and computational predictions based on genomic context. BioGRID is a curated repository that focuses exclusively on experimentally determined interactions, with each entry traceable to a specific publication and experimental system.

For a proteomics researcher, this difference has immediate practical consequences. When you upload your list of differentially expressed proteins from a mass spectrometry experiment, STRING will return a densely connected network with many predicted interactions, some of which may not have been directly observed in any single experiment. BioGRID will return a sparser network, but every edge in that network corresponds to a documented experimental finding.

The choice between these databases is not about which one is better in absolute terms. It is about matching the database to the question you are asking. If you are exploring the functional landscape of an uncharacterized protein set and want to generate hypotheses about biological processes, STRING provides the breadth you need. If you are testing whether a specific protein complex exists or whether two proteins physically associate under defined conditions, BioGRID gives you the experimental evidence required to support or refute that claim.

Consider how recent proteomics studies have used these tools. A 2025 study of dental pulp stem cells challenged with Porphyromonas gingivalis used STRINGDB alongside REACTOME to identify key proteins and networks involved in growth factor signalling, immune response, and wound healing [7]. The researchers needed functional enrichment and network context to interpret their comparative proteomic data, and STRING provided that layer of analysis. In contrast, a 2026 study mapping the human DBR1 interactome used immunopurification coupled to mass spectrometry and then compared their detected associations with BioGRID, finding that most of their newly identified interactions had not been previously reported [9]. That comparison against BioGRID served as a benchmark for novelty, because BioGRID represents the accumulated experimentally verified interaction knowledge.

## Data Content and Evidence Types

### STRING Evidence Channels

STRING organizes its evidence into several channels that feed into its confidence scoring system. These channels include neighbourhood, gene fusion, co-occurrence, co-expression, experimental data, curated databases, and text mining. The first three channels are computational predictions based on genomic context and are most relevant for prokaryotes and less-studied eukaryotes. Co-expression draws on transcriptomic data across many conditions. Experimental data and curated databases pull from primary interaction studies and resources such as BioGRID itself. Text mining scans the scientific literature for co-occurrence of gene names, which can generate both true associations and false positives.

For proteomics applications, the experimental and curated database channels carry the most weight, but the predicted channels can still be useful for generating hypotheses about proteins that lack direct interaction data. The key point is that STRING does not distinguish clearly between physical interactions and functional associations. A high STRING score may reflect co-expression or shared pathway membership instead of direct physical binding. This is acceptable for pathway enrichment and functional annotation, but it is insufficient for claims about physical interaction.

### BioGRID Curation Standards

BioGRID curates interactions from the primary literature, with each interaction record specifying the experimental system used, the publication source, and the interaction type. The database distinguishes between physical interactions and genetic interactions. Physical interactions include direct binding, co-purification, and other biochemical evidence. Genetic interactions include synthetic lethality, suppression, and other functional relationships inferred from genetic experiments.

The curation process is manual and systematic, which means BioGRID coverage reflects the historical focus of the molecular biology community. Model organisms such as Saccharomyces cerevisiae, Caenorhabditis elegans, Drosophila melanogaster, and Homo sapiens are well represented. Less-studied organisms may have sparse coverage, and interactions discovered in high-throughput screens may be listed alongside those from focused low-throughput studies without a built-in quality filter.

The DBR1 interactome study illustrates the value of BioGRID as a reference standard. The researchers mapped the human DBR1 interactome by immunopurification coupled to mass spectrometry and then compared their results with BioGRID to determine which associations were novel [9]. This workflow treats BioGRID as the accumulated experimentally verified knowledge base, and the comparison provides a measure of discovery novelty. If you are doing similar interactome mapping, this same comparison strategy can help you assess how much of your detected network is already known.

## Confidence Scoring and Interpretation

### STRING Combined Scores

STRING assigns each interaction a combined confidence score ranging from 0 to 1, with higher values indicating greater confidence that the interaction is real. The score integrates the individual evidence channels, with different weights applied to each channel. STRING also provides a medium confidence threshold of 0.400 and a high confidence threshold of 0.700, which are commonly used defaults in network analysis.

The practical implication for proteomics is that you must decide on a confidence threshold before you interpret the network. A low threshold will produce a dense network with many predicted edges, which can be useful for exploratory analysis but will include many false positives. A high threshold will produce a sparse network with fewer edges, which reduces false positives but may exclude genuine interactions that lack strong supporting evidence.

The choice of threshold depends on your downstream analysis. For gene ontology enrichment and pathway analysis, a medium confidence threshold is often appropriate because you are looking for functional themes instead of specific physical interactions. For hypothesis generation about specific protein complexes, a high confidence threshold is safer because you want to minimize the risk of pursuing a predicted interaction that does not hold up experimentally.

### BioGRID Evidence Quality

BioGRID does not provide a unified confidence score in the same way STRING does. Instead, each interaction record includes the experimental system and the publication source, allowing you to assess quality on a case-by-case basis. Interactions detected by affinity capture followed by mass spectrometry are common in proteomics contexts, and these are generally considered reliable evidence for physical association. Interactions detected by two-hybrid screens may include more false positives, particularly for membrane proteins and proteins that require post-translational modifications for binding.

When you use BioGRID data, you should examine the experimental system for the specific interactions that matter to your analysis. If you are building a network around a protein of interest and the key edges come from a single high-throughput screen, you may want to verify those interactions in the primary literature before drawing strong conclusions. If the edges come from multiple independent studies using different experimental systems, you can have greater confidence in those associations.

## Coverage and Organism Scope

### STRING Coverage Across Species

STRING covers thousands of organisms, including bacteria, archaea, and eukaryotes. The database uses orthology to transfer interaction information between species, which means that interactions known in one organism can be projected onto orthologous proteins in another organism. This orthology-based transfer is a major advantage for studying non-model organisms, where direct experimental interaction data may be scarce.

For proteomics researchers working on less-studied species, STRING is often the only practical option for network analysis. The orthology transfer allows you to leverage interaction knowledge from model organisms to interpret your data. However, you should be aware that orthology-based predictions are indirect evidence. An interaction predicted by orthology transfer may not exist in your organism of interest, even if it is well documented in the model organism.

### BioGRID Model Organism Focus

BioGRID has deep coverage for a smaller set of organisms, with a strong emphasis on yeast, worms, flies, and humans. The database also includes Arabidopsis thaliana, zebrafish, and several other species, but the depth of curation varies. For human proteomics, BioGRID is a valuable resource because it aggregates the extensive literature on human protein interactions.

The practical consequence is that BioGRID is most useful when you are working on a well-studied organism. If your proteomics experiment uses human cell lines, mouse tissues, or yeast cultures, BioGRID will provide substantial coverage of experimentally verified interactions. If your experiment uses a less common organism, BioGRID may return very few interactions for your protein list, and STRING will be the more useful resource.

## Usability and Workflow Integration

### STRING Web Interface and API

STRING provides a web interface that accepts a list of protein identifiers and returns a network visualization with functional enrichment analysis. The interface allows you to adjust confidence thresholds, select evidence channels, and export network images and data files. STRING also provides an API for programmatic access, which is useful for integrating network analysis into automated proteomics pipelines.

The web interface is straightforward for exploratory analysis. You paste your gene list, select your organism, and review the resulting network. The functional enrichment output includes gene ontology terms, KEGG pathways, and other annotations, which can be exported for further analysis. For reproducible workflows, the API allows you to script the analysis and document your parameters.

### BioGRID Search and Download Options

BioGRID provides a search interface that allows you to query by gene name, organism, or interaction type. The database also offers bulk downloads of the complete interaction dataset in multiple formats, including tab-delimited files and PSI-MI XML. These downloads are useful for building custom analysis pipelines or for performing systematic comparisons against your own interaction data.

For proteomics researchers, the BioGRID download files can be loaded into R or Python for custom analysis. The Bioconductor project provides packages for working with interaction data, and the broader bioinformatics ecosystem includes tools for network analysis that can consume BioGRID data [3]. If you prefer a graphical interface, the BioGRID website provides a straightforward search experience for looking up individual proteins or interaction sets.

### Integration with Proteomics Pipelines

Both databases can be integrated into proteomics analysis workflows, but the integration patterns differ. STRING is commonly used as a post-processing step after differential expression analysis, where the list of significant proteins is submitted for network and enrichment analysis. BioGRID is more commonly used as a reference database for validating detected interactions or for building custom interaction networks based on experimental evidence.

The Galaxy Training Network provides tutorials on proteomics analysis workflows that include steps for functional analysis and interaction mapping [4]. These tutorials emphasize reproducible analysis, which means documenting the database version, the parameters used, and the date of access. Similarly, the nf-core documentation describes community standards for reproducible bioinformatics pipelines, including the importance of version pinning for reference databases [5]. When you use STRING or BioGRID in your analysis, you should record the database version and access date so that your results can be reproduced or updated.

## At a Glance

| Feature | STRING | BioGRID |
|---------|--------|---------|
| Primary data type | Predicted and experimentally derived interactions, functional associations | Experimentally verified physical and genetic interactions |
| Evidence sources | Experimental data, curated databases, text mining, co-expression, genomic context predictions | Manual curation of primary literature |
| Confidence scoring | Combined score from 0 to 1 with adjustable thresholds | No unified score, evidence quality assessed per record |
| Organism coverage | Thousands of species with orthology transfer | Deep coverage for model organisms, limited for others |
| Best use case | Functional enrichment, pathway context, hypothesis generation for discovery proteomics | Validation of physical interactions, benchmarking novelty, mechanistic studies |
| Output format | Network visualization, enrichment tables, API access | Search results, bulk downloads in multiple formats |
| Reproducibility | Version and parameter dependent, requires documentation | Version dependent, requires documentation |

## Practical Workflow for Choosing Between STRING and BioGRID

### Step 1: Define Your Research Question

Before you open either database, write down the specific question you are trying to answer. Are you asking which biological processes are enriched in your differentially expressed protein list? Are you asking whether your protein of interest forms a complex with known partners? Are you asking how much of your detected interactome is novel compared with existing knowledge?

The answer to these questions determines the database choice. For process enrichment and functional context, STRING is the appropriate tool. For physical interaction validation and novelty assessment, BioGRID is the appropriate reference. For a comprehensive analysis, you may use both, with STRING providing the functional network and BioGRID providing the experimental evidence layer.

### Step 2: Prepare Your Protein List

Format your protein list according to the identifier requirements of each database. STRING accepts gene symbols, Ensembl identifiers, and several other formats. BioGRID accepts gene symbols and systematic identifiers depending on the organism. Check the identifier mapping before submission to avoid unnecessary errors.

For proteomics data, you should also decide whether to include all detected proteins or only those that passed your significance threshold. Including all detected proteins can be useful for understanding the full functional landscape, but it will also introduce noise from proteins that did not change between conditions. Including only significant proteins focuses the analysis on the biological response you are studying.

### Step 3: Run STRING for Functional Context

Submit your protein list to STRING and select your organism. Start with the default confidence settings and review the network that is generated. Examine the enrichment tables for gene ontology terms and pathways that are overrepresented in your list. These results will give you a functional interpretation of your proteomics data.

If the network is too dense to interpret, increase the confidence threshold. If the network is too sparse to be informative, decrease the threshold. Document the threshold you used, because it directly affects the results. Export the network image and the enrichment tables for your records.

### Step 4: Query BioGRID for Experimental Evidence

For the specific proteins that are central to your analysis, query BioGRID to see what experimental interaction evidence exists. Look at the experimental systems used to detect each interaction and the publications that reported them. This information tells you whether the interactions you are interested in have direct experimental support.

If you are mapping a novel interactome, download the BioGRID dataset for your organism and compare your detected interactions against the curated set. The DBR1 study used this approach to show that most of their detected associations were not previously reported in BioGRID [9]. This comparison provides a quantitative measure of novelty for your interaction data.

### Step 5: Integrate Results and Document Parameters

Combine the STRING functional analysis with the BioGRID experimental evidence to build a complete picture of your protein interaction landscape. Record the database versions, access dates, confidence thresholds, and identifier mappings used in your analysis. This documentation is essential for reproducibility, as emphasized in bioinformatics training resources from EMBL-EBI and the Galaxy Training Network [2][4].

## Options and Tradeoffs in Database Selection

### Using STRING Alone

For many proteomics studies, STRING alone provides sufficient analytical depth. The functional enrichment and network visualization capabilities are well suited for interpreting differential expression results. The confidence scoring system allows you to filter interactions according to your tolerance for false positives.

The tradeoff is that STRING does not distinguish clearly between physical interactions and functional associations. If you report that two proteins interact based on STRING analysis, you are making a claim that may not be supported by direct experimental evidence. For hypothesis generation, this is acceptable. For publication claims about physical interaction, it is not.

### Using BioGRID Alone

For mechanistic studies focused on known proteins, BioGRID alone may be sufficient. The experimentally verified interaction data provides a solid foundation for building interaction networks that can support claims about physical association. The genetic interaction data adds a functional dimension that is not available in STRING.

The tradeoff is that BioGRID coverage is limited for less-studied organisms and for proteins that have not been the subject of focused interaction studies. If your protein of interest has little BioGRID data, you will need to supplement with other resources or generate your own interaction data.

### Using Both Databases

The most robust approach is to use both databases in a complementary manner. STRING provides the broad functional network and enrichment analysis, while BioGRID provides the experimental evidence layer for specific interactions. This combined approach allows you to interpret your proteomics data in a functional context while also assessing the experimental support for individual interactions.

The 2025 dental pulp stem cell study used STRINGDB for network analysis of differentially expressed proteins, demonstrating the value of STRING for interpreting comparative proteomic data [7]. The 2026 DBR1 study used BioGRID as a comparison reference for assessing novelty of detected interactions [9]. Together, these studies illustrate the complementary roles of the two databases in proteomics research.

## Observations and Measurements in PPI Database Analysis

### Network Density and Confidence Thresholds

When you run STRING analysis on a proteomics dataset, the network density changes dramatically with the confidence threshold. At low thresholds, the network may include hundreds of edges connecting most of your proteins. At high thresholds, the network may fragment into small clusters or isolated nodes. These observations are informative about the overall connectivity of your protein set.

A highly connected network at high confidence suggests that your differentially expressed proteins participate in coordinated biological processes. A fragmented network suggests that your protein list includes diverse functional groups that do not share many interactions. Both patterns are biologically meaningful and should be interpreted in the context of your experimental system.

### Overlap Between Predicted and Experimentally Verified Interactions

If you run both STRING and BioGRID analysis on the same protein list, you can measure the overlap between predicted interactions and experimentally verified interactions. This overlap provides a sense of how much of the STRING network is supported by direct experimental evidence.

The DBR1 study found that most of their detected associations were not previously reported in BioGRID [9]. This observation indicates that mass spectrometry-based interactome mapping can discover novel interactions that are absent from curated databases. It also highlights the limitation of relying solely on existing databases for interpreting new interaction data.

### Functional Enrichment Consistency

When you perform gene ontology enrichment analysis on your protein list, the results should be consistent with the biological context of your experiment. If you are studying a stress response, you expect enrichment for stress-related terms. If you are studying a developmental process, you expect enrichment for developmental terms.

Inconsistencies between your expected biology and the enrichment results may indicate problems with your protein list, your identifier mapping, or your choice of background set. The muscle aging study used gene ontology analysis to define the aging matreotype and found that aging was characterized by structural alterations, synaptic transmission changes, and a decline in ECM and angiogenesis-related biological processes [10]. This type of functional interpretation depends on reliable enrichment analysis, which in turn depends on appropriate database selection.

## Records and Documentation for Reproducible Analysis

### Database Version Tracking

Both STRING and BioGRID release updated versions on a regular basis. The content changes between versions, with new interactions added and existing records corrected. If you do not record the version you used, your analysis cannot be reproduced exactly.

Record the database version, the access date, and the URL used for each analysis. This information should be included in your laboratory notebook and in the methods section of any publication that reports the analysis. The nf-core documentation emphasizes the importance of version pinning for reproducible workflows [5], and the same principle applies to database usage.

### Parameter Documentation

For STRING analysis, record the confidence threshold, the evidence channels included, and the organism selection. For BioGRID analysis, record the search parameters, the interaction types included, and the download file version. These parameters directly affect the results and must be documented for reproducibility.

The Galaxy Training Network provides tutorials that emphasize reproducible analysis workflows, including the documentation of parameters and data sources [4]. Following these practices ensures that your PPI database analysis can be reviewed, validated, and updated as new data become available.

### Identifier Mapping Records

Proteomics data often uses protein identifiers from UniProt or Ensembl, while PPI databases may use gene symbols or other identifiers. The mapping between these identifier systems can introduce errors if not performed carefully. Record the identifier mapping method and the version of the mapping resource used.

The NCBI provides a range of data resources and search systems that can be used for identifier mapping and cross-referencing [1]. Using official mapping resources reduces the risk of identifier errors in your PPI analysis.

## Common Failure Patterns in PPI Database Analysis

### Overinterpreting Predicted Interactions

The most common failure in STRING analysis is treating predicted interactions as experimentally verified facts. STRING includes many interactions that are based on text mining, co-expression, or genomic context predictions. These interactions may be useful for generating hypotheses, but they are not evidence of physical binding.

To avoid this failure, always check the evidence channels supporting the interactions that are central to your conclusions. If the key interactions are supported only by text mining or co-expression, you should verify them experimentally before making strong claims.

### Ignoring Organism Specificity

Another common failure is using interaction data from one organism to draw conclusions about another organism without considering orthology relationships. STRING transfers interactions between organisms based on orthology, but this transfer is an inference, not a direct observation. An interaction that is well documented in yeast may not exist in humans, even if the orthologous proteins are similar.

To avoid this failure, always check the organism context of the interactions you are using. If you are studying human proteins, prioritize interactions that have been directly observed in human cells or with human proteins.

### Neglecting Version Differences

A third failure pattern is neglecting to record the database version, making it impossible to reproduce the analysis or to understand why results differ between analyses performed at different times. Databases are updated regularly, and the interaction content changes with each release.

To avoid this failure, adopt a documentation practice that includes database versions and access dates for every analysis. This practice is standard in reproducible bioinformatics workflows and is emphasized in training resources from EMBL-EBI and the Carpentries [2][6].

### Using Inappropriate Confidence Thresholds

A fourth failure pattern is using a confidence threshold that does not match the research question. A low threshold produces a dense network with many false positives, while a high threshold produces a sparse network that may miss genuine interactions. The appropriate threshold depends on whether you are doing exploratory analysis or testing specific hypotheses.

To avoid this failure, test multiple thresholds and examine how the network structure changes. Choose the threshold that provides the most informative network for your specific question, and document the rationale for your choice.

## Limitations and Interpretation Boundaries

### STRING Limitations

STRING does not distinguish between physical interactions and functional associations. A high confidence score may reflect co-expression or shared pathway membership instead of direct physical binding. This limitation is acceptable for functional enrichment analysis but is critical for claims about physical interaction.

STRING also includes text mining evidence, which can introduce false positives from literature co-occurrence. Gene names that frequently appear together in publications may be assigned high scores even if the proteins do not interact. This is particularly problematic for well-studied proteins that appear in many publications.

### BioGRID Limitations

BioGRID coverage reflects the historical focus of the molecular biology community. Proteins that have been studied extensively have deep interaction coverage, while less-studied proteins may have sparse or absent interaction data. This coverage bias means that the absence of an interaction in BioGRID does not mean the interaction does not exist.

BioGRID also includes interactions from high-throughput screens that may have higher false positive rates than focused low-throughput studies. The database does not provide a unified quality score, so you must assess evidence quality on a case-by-case basis by examining the experimental system and the publication source.

### General Limitations of PPI Databases

Both databases are incomplete representations of the true interactome. Many genuine interactions have not been detected or curated, and some curated interactions may be false positives. The DBR1 study demonstrated that mass spectrometry-based interactome mapping can identify many associations that are absent from BioGRID [9], indicating that current databases do not capture the full extent of protein interactions.

The denaturation-based assay study showed that different experimental workflows capture different PPI networks, with protein size, structural complexity, hydrophobicity, and localization influencing detection in a workflow-specific manner [8]. This finding underscores the importance of understanding the experimental context of interaction data and the limitations of any single detection method.

## Safety and Regulatory Context for PPI Database Use

### Data Integrity and Reproducibility Requirements

In regulated research environments, the use of PPI databases must meet data integrity standards that support reproducibility and auditability. This means documenting database versions, access dates, analysis parameters, and identifier mappings. The nf-core documentation provides community standards for reproducible bioinformatics pipelines that can be adapted to PPI analysis [5].

For studies that support regulatory submissions or clinical decisions, the analytical methods must be validated and documented according to applicable standards. The choice of PPI database and the parameters used for analysis should be justified in the study documentation.

### Professional Escalation Criteria

If your PPI analysis produces results that are inconsistent with established biology or that cannot be reproduced with different database versions, you should escalate the issue to a bioinformatics specialist or a colleague with expertise in interaction databases. Similarly, if you are uncertain about the interpretation of specific interactions or the appropriate confidence threshold for your analysis, seek guidance before drawing conclusions.

The EMBL-EBI training resources provide pathways for developing bioinformatics skills, including the use of interaction databases [2]. The Carpentries lessons provide foundational computing skills that support reproducible analysis [6]. Investing in these skills reduces the risk of analytical errors and improves the quality of your PPI analysis.

## Building a Reproducible PPI Database Decision Log for Your Proteomics Project

A recurring problem in proteomics analysis is that researchers make an implicit choice between STRING and BioGRID without documenting why they chose one database over the other or how that choice shaped their biological conclusions. This lack of documentation becomes a serious issue when you need to revisit the analysis months later, when a collaborator asks how you arrived at a particular network interpretation, or when a reviewer requests the exact parameters that produced your enrichment results. A structured decision log solves this problem by forcing you to record the rationale, parameters, and expected outputs of each database query before you run it. This section provides a practical framework for building that log, with specific fields, example entries, and common failure patterns to avoid.

### Why a Decision Log Matters for PPI Analysis

The choice between STRING and BioGRID is not a one-time decision that applies to your entire project. You will likely use both databases at different stages of your analysis, and the order in which you use them affects how you interpret your results. If you run STRING first and generate a dense functional network, you may unconsciously bias your subsequent BioGRID queries toward interactions that appear in that network. If you run BioGRID first and find sparse experimental coverage for your proteins, you may conclude that your protein list is poorly characterized when in fact the issue is limited curation for your organism.

A decision log addresses this problem by separating the analytical steps from the interpretation. Each entry records what you queried, why you queried it, and what you expected to find. When you later compare results across databases, you can trace each conclusion back to a specific query with documented parameters. This practice aligns with the reproducibility standards emphasized in bioinformatics training resources from EMBL-EBI and the Galaxy Training Network, which both stress the importance of documenting data sources and analysis parameters [2][4].

### Core Fields for Each Decision Log Entry

Create one log entry for each distinct database query you perform. Each entry should contain the following fields.

**Query identifier and date.** Assign a unique identifier such as PPI-001 and record the date. This allows you to reference specific queries in your laboratory notebook and in the methods section of publications.

**Research question addressed.** Write one sentence describing the specific question this query answers. Examples include identifying enriched biological processes in the differentially expressed protein list, determining whether protein X has experimentally verified physical interactions with any member of the candidate complex, or assessing how many of the detected interactions in the mass spectrometry dataset are already present in curated databases.

**Database and version.** Record whether you used STRING or BioGRID and the exact version number. Both databases release updates regularly, and interaction content changes between versions. The nf-core documentation emphasizes version pinning for reproducible workflows, and the same principle applies to database queries [5].

**Input identifier list.** Record the exact list of identifiers you submitted, including the identifier type such as UniProt accession, Ensembl gene ID, or gene symbol. Also record the mapping method if you converted identifiers from one system to another. The NCBI provides official search systems and data resources that can be used for identifier mapping and cross-referencing [1].

**Organism selection.** Record the organism you selected in the database interface. This is critical because both databases use organism-specific data, and selecting the wrong organism will produce misleading results.

**Key parameters.** For STRING, record the confidence threshold, the evidence channels you included or excluded, and whether you used the default settings. For BioGRID, record the interaction types you included such as physical, genetic, or both, and any filters you applied.

**Expected output.** Write one sentence describing what you expected the query to return. This forces you to articulate your hypothesis before running the analysis and makes it easier to identify unexpected results.

**Actual output and interpretation.** After running the query, record what the database returned and how you interpreted it. Note any discrepancies between expected and actual output.

### Example Decision Log Entries

Consider a proteomics experiment that identifies 150 differentially expressed proteins in a human cell line treated with a stress agent. Your analysis might produce the following log entries.

**PPI-001.** Date 2026-03-15. Research question: Which biological processes are enriched in the 150 differentially expressed proteins? Database: STRING version 12.0. Input: 150 UniProt accessions. Organism: Homo sapiens. Parameters: medium confidence threshold 0.400, all evidence channels enabled. Expected output: enrichment for stress response and apoptosis terms. Actual output: significant enrichment for response to unfolded protein, apoptotic signaling pathway, and mRNA splicing. Interpretation: the protein list captures both the expected stress response and an unexpected splicing component that warrants further investigation.

**PPI-002.** Date 2026-03-15. Research question: Do any of the 150 differentially expressed proteins have experimentally verified physical interactions with HSP90AB1? Database: BioGRID version 4.4.235. Input: HSP90AB1 gene symbol. Organism: Homo sapiens. Parameters: physical interactions only. Expected output: a list of known HSP90AB1 interaction partners. Actual output: 87 physical interaction partners, of which 12 appear in the differentially expressed protein list. Interpretation: the overlap between the stress response proteome and the HSP90AB1 interactome supports a role for this chaperone in the cellular response to the stress agent.

**PPI-003.** Date 2026-03-16. Research question: How many of the protein-protein interactions detected in our immunoprecipitation mass spectrometry experiment are already present in curated databases? Database: BioGRID version 4.4.235. Input: list of 45 detected interaction pairs. Organism: Homo sapiens. Parameters: physical interactions only. Expected output: most detected interactions would be present in BioGRID. Actual output: only 8 of 45 detected pairs were present in BioGRID. Interpretation: the experiment identified a substantial number of novel interactions, consistent with the finding from the DBR1 interactome study that mass spectrometry can detect associations not previously reported in BioGRID [9].

### Using the Decision Log to Compare Databases

The decision log becomes most valuable when you use it to compare results across databases for the same protein list. After you have completed STRING and BioGRID queries for your differentially expressed proteins, create a comparison table that lists the following for each protein or protein pair: the STRING confidence score, the evidence channels supporting the STRING prediction, the BioGRID experimental systems and publications, and whether the interaction appears in both databases.

This comparison reveals the degree of overlap between predicted and experimentally verified interactions. A protein pair with a high STRING score and strong BioGRID evidence has robust support. A protein pair with a high STRING score but no BioGRID entry may be a genuine interaction that has not been experimentally characterized, or it may be a false positive driven by text mining or co-expression evidence. The DBR1 study demonstrated that many mass spectrometry-detected associations are absent from BioGRID [9], which means you should not interpret absence from BioGRID as evidence against an interaction.

### Common Failure Patterns in PPI Database Documentation

**Recording only the database name without the version.** This is the most common documentation failure. If you record only STRING or BioGRID without the version number, you cannot reproduce the analysis after the database updates. Always record the full version string.

**Failing to record the identifier mapping method.** Proteomics data often uses UniProt accessions while PPI databases may expect gene symbols. If you convert identifiers using a mapping tool, record which tool and which version you used. Different mapping tools can produce different results, and the NCBI provides official resources for cross-referencing identifiers [1].

**Not recording the confidence threshold rationale.** The choice between medium and high confidence thresholds in STRING directly affects network density and enrichment results. Record also the threshold value but also the reason you chose it. For exploratory analysis, a medium threshold is often appropriate. For hypothesis testing about specific complexes, a high threshold reduces false positives.

**Treating the decision log as an afterthought.** If you create the decision log after completing the analysis, you will likely forget key parameters or reconstruct them inaccurately. Create the log entries before running each query and update them immediately after.

### Professional Escalation Criteria

If your decision log reveals inconsistencies that you cannot resolve, escalate the issue to a colleague with bioinformatics expertise. Specific situations that warrant escalation include: STRING and BioGRID return contradictory results for the same protein pair, the enrichment results are inconsistent with the known biology of your experimental system, or the overlap between your detected interactions and curated databases is unexpectedly low or high. The EMBL-EBI training resources provide pathways for developing the skills needed to troubleshoot these issues [2], and the Carpentries lessons offer foundational computing skills that support reproducible analysis [6].

### Integrating the Decision Log into Your Laboratory Records

The decision log should be maintained as a spreadsheet or a plain text file in your laboratory information management system. Each entry should be timestamped and linked to the raw data files from your proteomics experiment. When you publish your results, include the decision log as supplementary material or describe its contents in the methods section. This practice supports the reproducibility standards emphasized in the nf-core documentation [5] and the Galaxy Training Network tutorials [4], and it ensures that your PPI database analysis can be reviewed, validated, and updated as new data become available.

## Frequently Asked Questions

### What is the main difference between STRING and BioGRID?

STRING aggregates multiple evidence types, including predictions, text mining, and experimental data, to provide a broad functional network with confidence scores. BioGRID curates experimentally verified interactions from the primary literature, providing a reference set of physical and genetic interactions. STRING is better for functional enrichment and hypothesis generation, while BioGRID is better for validating specific interactions and assessing novelty.

### Which database should I use for functional enrichment analysis of my proteomics data?

Use STRING for functional enrichment analysis. STRING provides gene ontology enrichment and pathway analysis alongside network visualization, which is well suited for interpreting lists of differentially expressed proteins. The 2025 dental pulp stem cell study used STRINGDB for deeper analysis of differentially expressed proteins [7], and the alopecia areata study used STRING for protein-protein interaction network analysis [11].

### How do I choose a confidence threshold in STRING?

The confidence threshold depends on your research question. For exploratory analysis and functional enrichment, a medium confidence threshold of 0.400 is often appropriate. For hypothesis generation about specific protein complexes, a high confidence threshold of 0.700 reduces false positives. Test multiple thresholds and examine how the network structure changes before selecting a value for your final analysis.

### Can I use BioGRID to assess whether my detected interactions are novel?

Yes. Download the BioGRID dataset for your organism and compare your detected interactions against the curated set. The DBR1 study used this approach to show that most of their detected associations were not previously reported in BioGRID [9]. This comparison provides a quantitative measure of novelty for your interaction data.

### Does BioGRID include predicted interactions?

No. BioGRID curates experimentally verified interactions from the primary literature. Each interaction record specifies the experimental system used and the publication source. If you need predicted interactions or functional associations, use STRING instead.

### How do I document my PPI database analysis for reproducibility?

Record the database version, access date, confidence threshold, evidence channels, organism selection, and identifier mapping method for each analysis. Include this information in your laboratory notebook and in the methods section of any publication. The nf-core documentation and the Galaxy Training Network provide guidance on reproducible bioinformatics workflows [5][4].

### What should I do if my protein list returns very few interactions in BioGRID?

If your organism or proteins have sparse BioGRID coverage, use STRING to obtain a broader network based on orthology transfer and multiple evidence channels. You can also check whether your identifier mapping is correct and whether you are using the appropriate organism selection. For less-studied organisms, STRING is often the only practical option for network analysis.

### How do I avoid overinterpreting predicted interactions from STRING?

Check the evidence channels supporting the interactions that are central to your conclusions. If the key interactions are supported only by text mining or co-expression, verify them experimentally before making strong claims. STRING does not distinguish between physical interactions and functional associations, so you should not treat predicted interactions as evidence of physical binding.

## Related Bioinformatics Guides

- [Multi-Omics Data Integration: A Comparative Framework for Choosing the Right Method](/knowledge/bioinformatics/multi-omics-data-integration-a-comparative-framework-for-choosing-the-right-method)
- [STRING Database and Protein-Protein Interaction Networks](/knowledge/bioinformatics/string-database-and-protein-protein-interaction-networks)
- [Genomic Data Analysis Tools: A Comparative Guide for Researchers](/knowledge/bioinformatics/genomic-data-analysis-tools-a-comparative-guide-for-researchers)
- [Spatial Proteomics vs. Single-Cell Proteomics: Choosing the Right Approach](/knowledge/bioinformatics/spatial-proteomics-vs-single-cell-proteomics-choosing-the-right-approach)
- [Proteomics Mass Spectrometry: From Sample Preparation to Data Analysis](/knowledge/bioinformatics/proteomics-mass-spectrometry-from-sample-preparation-to-data-analysis)

## References and Further Reading

- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Bioconductor](https://bioconductor.org/). Bioconductor Project.
- [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.
- [nf-core Documentation](https://nf-co.re/docs). nf-core.
- [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.
- [Comparative proteomic analysis of Porphyromonas gingivalis challenged human dental pulp stem cells.](https://doi.org/10.1038/s41598-025-29191-z). 2025.
- [Exploring How Workflow Variations in Denaturation-Based Assays Impact Global Protein-Protein Interaction Predictions.](https://doi.org/10.1016/j.mcpro.2025.101479). 2026.
- [The human DBR1 interactome reveals coupling between intron lariat turnover, pre-mRNA splicing, and RNA quality control pathways.](https://doi.org/10.1016/j.cstres.2026.100178). 2026.
- [Transcriptional profiling of male mouse muscle across aging stages: a gene ontology analysis of the muscle matreotype.](https://doi.org/10.1186/s12864-026-12677-z). 2026.
- [Differential proteomics of lesional vs. non-lesional biopsies revealed non-immune mechanisms of alopecia areata](https://doi.org/10.1038/s41598-017-18282-1). Scientific Reports, 2018.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.