# STRING Analysis in Proteomics: How to Use Protein-Protein Interaction Networks to Interpret Your Results

Proteomics experiments generate long lists of differentially abundant proteins, but a list alone does not explain the biology. STRING (Search Tool for the Retrieval of Interacting Genes/Proteins) addresses this problem by assembling protein-protein interaction networks from multiple evidence channels and by providing functional enrichment analysis for any set of proteins. This article explains how to prepare proteomics data for STRING, how to build and interpret interaction networks, how to evaluate network quality, and how to export results for publication. The workflow applies to label-free quantification, tandem mass tag (TMT), data-independent acquisition (DIA), and other quantitative mass spectrometry outputs.

## What STRING Does and Where It Fits in a Proteomics Workflow

STRING systematically collects and integrates protein-protein interactions, covering both physical interactions and functional associations. The data come from automated text mining of the scientific literature, computational interaction predictions from co-expression, conserved genomic context, databases of interaction experiments, and known complexes and pathways from curated sources. All interactions are critically assessed, scored, and automatically transferred to less well-studied organisms using hierarchical orthology information. The database can be accessed via the website, programmatically, and through bulk downloads.

In a typical proteomics workflow, STRING is used after differential abundance analysis. The input is a list of proteins that changed between experimental conditions, and the output is a network that shows which of those proteins are known or predicted to work together. This network interpretation helps researchers move from a list of names to a testable hypothesis about the underlying mechanism.

The most recent developments in STRING version 12.0 include the ability to create, browse, and analyze a full interaction network for any novel genome of interest by submitting its complement of encoded proteins. The co-expression channel now uses variational auto-encoders to predict interactions and covers two new sources: single-cell RNA-seq and experimental proteomics data. The confidence in each experimentally derived interaction is estimated based on the detection method used and communicated to the user in the web interface. Functional enrichment analysis is fully available for user-submitted genomes.

## At a Glance: STRING Inputs, Outputs, and Decision Points

| Workflow Step | What You Provide | What STRING Returns | Key Decision You Make |
| --- | --- | --- | --- |
| Organism selection | Taxonomic identifier or common name | Organism-specific interaction database | Confirm the correct species, especially for non-model organisms |
| Protein list input | Gene symbols, UniProt IDs, or protein sequences | Mapped identifiers with match status | Check unmapped entries and resolve identifier mismatches |
| Network construction | Confidence threshold and network type | Edges with combined scores and evidence channels | Choose a threshold that balances sensitivity and specificity |
| Enrichment analysis | Background set and query set | Functional terms with false discovery rate values | Select a background set that matches your experimental scope |
| Network visualization | Clustering method and display options | Clustered network with colored nodes and edges | Adjust clustering to reveal biologically meaningful modules |
| Export | Image format and data format | Publication-ready figures and tabular edge lists | Verify that exported files meet journal requirements |

## Preparing Proteomics Data for STRING Input

### Identifier Selection and Mapping

The first decision is which identifier type to use. STRING accepts gene symbols, Ensembl identifiers, UniProt identifiers, and several other formats. For mass spectrometry data, UniProt accession numbers are often the native output from database search engines such as MaxQuant, Proteome Discoverer, or Spectronaut. Gene symbols are frequently easier to interpret and are the default display format in STRING.

Before submitting your list, check whether your identifiers are current. Protein databases are updated regularly, and obsolete accession numbers may fail to map. The [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/) provide official descriptions of databases, search systems, sequence resources, and analysis services, which can help you verify identifier formats and retrieve current annotations. If you are working with a species that has multiple annotation versions, confirm that your identifiers match the version used by STRING.

### Handling Multiple Isoforms and Redundant Entries

Mass spectrometry often detects multiple isoforms or redundant entries for the same gene. STRING maps proteins to genes, so multiple isoforms of the same gene collapse into a single network node. This is usually desirable because the interaction evidence is gene-centric. However, if your analysis depends on isoform-specific functions, you should note that STRING does not distinguish between isoforms in its network display.

For quantitative proteomics, you should decide whether to aggregate abundance values across isoforms before submitting the list to STRING. If you submit a list of differentially abundant proteins, the aggregation step does not affect the network output, but it does affect the number of entries in your input list. Keep a record of how many unique genes your protein list represents, because this number is the denominator for enrichment statistics.

### Handling Unmapped Entries

Some proteins in your list will not map to STRING. This happens for several reasons: the identifier is obsolete, the protein is from a contaminant database, or the organism is not well represented in STRING. Record the number of unmapped entries and the reasons for mapping failure. If more than a small fraction of your list fails to map, check your identifier format before proceeding. The STRING website reports mapping statistics for each submitted list, and you should inspect these statistics before interpreting the network.

### Background Set Selection for Enrichment Analysis

Functional enrichment analysis compares your query list against a background set. The default background in STRING is the entire genome of the selected organism. This is appropriate when your query list comes from a discovery experiment that measured the whole proteome. However, if your mass spectrometry experiment only detected a subset of the proteome, the detected protein set is a more appropriate background. Using the detected set as background reduces bias from proteins that are not measurable by your platform.

For targeted proteomics approaches such as OLINK, the background set should be the panel of measured proteins, not the entire genome. A [study of plasma proteomic biomarkers in the UK Biobank](https://pubmed.ncbi.nlm.nih.gov/38710337) used OLINK proteomics with 1,463 plasma proteins and performed pathway analysis on the measured protein set. This approach avoids the inflation that occurs when the background includes proteins that could not be detected by the assay.

## Building a Protein-Protein Interaction Network in STRING

### Organism Selection

Select the organism that matches your experimental system. STRING uses taxonomic identifiers, and the web interface provides a search box for common names. For human, mouse, rat, zebrafish, and other common model organisms, the selection is straightforward. For non-model organisms, confirm that your species is present and that the genome annotation is current.

If your organism is not in STRING, version 12.0 allows you to submit the complement of encoded proteins for a novel genome of interest. This feature enables network analysis for any sequenced genome, even if the organism has no precomputed interactions. The interaction predictions are transferred from well-studied organisms using hierarchical orthology information.

### Input Formats and List Size

STRING accepts a pasted list of identifiers, a file upload, or a query to the application programming interface. The web interface handles lists of several thousand proteins, but very large lists produce dense networks that are difficult to interpret. For a typical differential proteomics experiment with hundreds of differentially abundant proteins, the full list is appropriate for enrichment analysis. For network visualization, you may want to filter to the most significant proteins or use the network clustering features to reduce visual complexity.

### Confidence Threshold Selection

STRING assigns a confidence score to each interaction, and the user selects a threshold for network display. The default is medium confidence, which balances sensitivity and specificity. High confidence produces a sparser network with fewer false-positive edges, while low confidence produces a denser network with more predicted interactions.

The choice of threshold depends on your research question. If you are looking for well-established interactions to support a mechanism, high confidence is appropriate. If you are exploring novel connections and want to generate hypotheses, medium or low confidence may be more useful. Document the threshold you used, because it affects the network topology and the interpretation of clusters.

### Network Types: Full Network versus Physical Subnetwork

STRING distinguishes between full networks and physical subnetworks. The full network includes all interaction types: physical binding, co-expression, genetic interactions, and pathway associations. The physical subnetwork includes only direct physical interactions. For proteomics data, the full network is usually more informative because functional associations can reveal pathway relationships even when the proteins do not physically bind.

The evidence channels are displayed as colored edges in the network view. Each edge color corresponds to a different evidence type: text mining, experiments, databases, co-expression, neighborhood, gene fusion, and co-occurrence. Inspecting the edge colors helps you understand why two proteins are connected. A connection supported by experimental evidence is stronger than a connection supported only by text mining.

## Interpreting Network Clusters and Functional Modules

### Cluster Identification

Proteins that participate in the same biological process often cluster together in the network. STRING provides clustering options that group densely connected nodes. The clustering algorithm can be adjusted, and the resulting clusters can be examined for functional coherence.

When you identify a cluster, examine the proteins within it and ask whether they share a common function. For example, a cluster of complement and coagulation proteins would suggest that these pathways are coordinately regulated in your experiment. A [study of serum proteomics in thymoma patients](https://pubmed.ncbi.nlm.nih.gov/36991043) used STRING to assess the interaction of differential proteins and found that the complement and coagulation cascade was enriched, with von Willebrand factor, coagulation factor V, and vitamin K-dependent protein C upregulated.

### Hub Proteins and Network Centrality

Hub proteins are nodes with many connections. In a proteomics network, hubs may represent central regulators or abundant proteins that interact with many partners. Interpret hubs with caution: some hubs are genuinely important regulators, while others are highly abundant proteins that appear in many interaction records.

The STRING network view does not compute centrality metrics directly, but you can export the network and analyze it with other tools. If you identify a hub that is biologically plausible, consider validating its role with orthogonal experiments such as Western blot or targeted assays.

### Integrating Transcriptomics and Proteomics

Many experiments generate both transcriptomics and proteomics data. STRING can be used to analyze each data set separately and then compare the enriched pathways. A [study of lens-induced myopia in mice](https://pubmed.ncbi.nlm.nih.gov/37819745) used RNA-sequencing and data-independent acquisition liquid chromatography tandem mass spectrometry, then used Gene Ontology, KEGG annotation, and STRING databases to identify significantly affected pathways in both data sets. The transcriptomic and proteomic data showed low correlation at the individual gene level, but the pathway-level analysis revealed that the two data sets complemented one another.

This observation has a practical implication: do not expect mRNA and protein abundance to correlate strongly for every gene. Instead, compare the enriched pathways and interaction networks from each data set. Pathways that are enriched in both data sets are more likely to be biologically meaningful.

## Functional Enrichment Analysis with STRING

### Gene Ontology Enrichment

STRING performs enrichment analysis against Gene Ontology categories: biological process, molecular function, and cellular component. The output includes the number of query proteins in each term, the expected number by chance, and the false discovery rate.

For proteomics data, the cellular component category is often informative because it reveals where the differentially abundant proteins are located. The [thymoma serum proteomics study](https://pubmed.ncbi.nlm.nih.gov/36991043) found that the differential proteins were primarily exocrine and serum membrane proteins involved in controlling immunological responses and antigen binding. This localization information helps interpret the biological context of the changes.

### KEGG Pathway Enrichment

KEGG pathway enrichment identifies metabolic and signaling pathways that are overrepresented in your protein list. The same [thymoma study](https://pubmed.ncbi.nlm.nih.gov/36991043) found that the differential proteins play a significant role in the complement and coagulation cascade and the phosphoinositide 3-kinase (PI3K)/protein kinase B (AKT) signal pathway.

When interpreting KEGG results, check the number of proteins from your list that map to each pathway. A pathway with three proteins out of fifty is less convincing than a pathway with twenty proteins out of fifty. The enrichment p-value accounts for this, but you should also inspect the actual proteins to confirm that they are functionally related.

### Other Enrichment Categories

STRING also provides enrichment for protein domains, tissue expression, and subcellular localization. These categories can be useful for proteomics data because they connect your protein list to features that are not captured by pathway annotations. For example, enrichment for a specific protein domain might suggest a common binding mechanism, and tissue enrichment might indicate that your differentially abundant proteins are derived from a particular cell type.

### Multiple Testing Correction

Enrichment analysis tests many terms simultaneously, so multiple testing correction is essential. STRING reports false discovery rate values, and you should use these corrected values instead of raw p-values when selecting significant terms. A common threshold is a false discovery rate below 0.05, but you may choose a more stringent threshold for large protein lists.

## Exporting STRING Results for Publication

### Image Export Options

STRING provides several image export formats, including PNG and SVG. For publication, vector formats such as SVG are preferred because they scale without losing resolution. Before exporting, adjust the network layout, node size, and edge thickness to produce a readable figure.

The exported image should include a legend that explains the edge colors and the confidence score. STRING includes this information in the image, but you should verify that the legend is legible at the final figure size. Journals often require a minimum font size for figure text, so check the exported image at the dimensions you plan to submit.

### Tabular Export of Network Data

In addition to images, STRING exports tabular data that lists the edges with their confidence scores and evidence channels. This table is useful for supplementary materials and for downstream analysis. The table includes the two protein identifiers, the combined score, and the individual evidence scores.

Keep the tabular export as a permanent record of your analysis. If you need to reproduce the network later or if a reviewer asks for the underlying data, the table provides the complete information.

### Reproducibility Considerations

Reproducibility requires documenting the STRING version, the organism, the input identifiers, the confidence threshold, and the date of analysis. STRING is updated regularly, and interactions are added or removed with each version. A network generated with version 11.5 may differ from the same query in version 12.0.

The [Galaxy Training Network](https://training.galaxyproject.org/) provides accessible workflow training and analysis tutorials that emphasize reproducibility in bioinformatics workflows. Following similar practices for your STRING analysis ensures that other researchers can reproduce your network and enrichment results.

## Programmatic Access to STRING

### Application Programming Interface

STRING provides an application programming interface (API) that allows programmatic access to the database. The API supports identifier mapping, network retrieval, and enrichment analysis. For researchers who analyze many protein lists or who want to integrate STRING into a larger workflow, the API is more efficient than the web interface.

The API requires a format for the query and returns results in a machine-readable format. The STRING website documents the available endpoints and parameters. Before writing a script, test the API with a small query to confirm that you are using the correct parameter names and values.

### Bulk Downloads

STRING offers bulk downloads of the complete interaction data for each organism. These downloads are large and are intended for researchers who need to analyze the full interaction network offline. If you only need interactions for a few hundred proteins, the web interface or API is more practical.

Bulk downloads are useful for building custom analysis pipelines that combine STRING data with other resources. For example, [ProteomicsDB](https://pubmed.ncbi.nlm.nih.gov/29106664) integrates protein-protein interaction information from STRING with quantitative mass spectrometry data, transcriptomics data from NCBI GEO, functional annotations from KEGG, and drug sensitivity data. This integration enables protein-centric exploration of multiple data sources.

## Common Failure Patterns and How to Avoid Them

### Identifier Mapping Failures

The most common failure in STRING analysis is poor identifier mapping. This occurs when the input list uses a format that STRING does not recognize or when the identifiers are outdated. To avoid this problem, check the mapping statistics on the STRING results page. If more than 10 percent of your list fails to map, investigate the cause before proceeding.

For mass spectrometry data, contaminant proteins are a frequent source of unmapped entries. Many search engines append a contaminant prefix to common laboratory contaminants such as keratins and trypsin. Remove these entries from your list before submitting to STRING.

### Overinterpretation of Enrichment Results

Enrichment analysis identifies terms that are overrepresented in your list, but overrepresentation does not prove biological relevance. A term can be enriched because of a few highly connected proteins or because of technical artifacts. Before concluding that a pathway is important, inspect the individual proteins that contribute to the enrichment.

The [UK Biobank dementia study](https://pubmed.ncbi.nlm.nih.gov/38710337) illustrates the value of careful interpretation. The study identified 11 plasma proteins with strong mediating effects between cardiovascular health and dementia risk, with GDF15 having the strongest association. A first principal component with 10 top mediators mediated 53.6 percent of the effect. This analysis went beyond simple enrichment to quantify the contribution of specific proteins.

### Ignoring the Background Set

Using the wrong background set is a common error. If you use the whole genome as background when your experiment only measured a subset of proteins, the enrichment results will be biased toward the measurable proteins. This bias can produce false-positive pathway enrichment.

For mass spectrometry experiments, the background set should be all proteins that were quantified in the experiment, beyond the differentially abundant proteins. Most proteomics software can export the full list of quantified proteins, and you can use this list as the STRING background.

### Confusing Correlation with Causation

A protein-protein interaction network shows associations, not causal relationships. Two proteins that are connected in STRING may be co-regulated without directly interacting. The network is a hypothesis-generating tool, not a proof of mechanism.

When you identify a network module that is relevant to your research question, design follow-up experiments to test the relationship. For example, if a cluster of immune-related proteins is enriched in your data, validate the key proteins with Western blot or targeted assays before claiming a mechanistic role.

## Limitations of STRING for Proteomics Data

### Coverage Gaps in Interaction Databases

STRING coverage depends on the available interaction data for each organism. Well-studied organisms such as human and mouse have extensive interaction records, while less-studied organisms have sparse coverage. For non-model organisms, the network may be dominated by interactions transferred from model organisms via orthology, and these transferred interactions are less reliable than direct experimental evidence.

The [STRING database in 2023 article](https://pubmed.ncbi.nlm.nih.gov/36370105) notes that novel interactions continue to be discovered and that information remains scattered across different database resources, experimental modalities, and levels of mechanistic detail. This means that your network may miss recently discovered interactions that have not yet been integrated into STRING.

### Bias Toward Well-Studied Proteins

Interaction databases are biased toward well-studied proteins. Highly abundant proteins and proteins that have been investigated for decades have more interaction records than poorly characterized proteins. This bias can make well-studied proteins appear as hubs in the network, even when their biological importance in your specific experiment is unclear.

When interpreting hub proteins, consider whether the hub status reflects genuine biological centrality or simply the amount of published research on that protein. The confidence score for each interaction provides some guidance, but it does not correct for the overall bias in literature coverage.

### Co-Expression Predictions from Proteomics Data

The co-expression channel in STRING now includes experimental proteomics data as a source. This is a recent development, and the predictions from this channel should be interpreted with the same caution as other computational predictions. Co-expression does not prove physical interaction, and the correlation may reflect shared regulation instead of direct binding.

The use of variational auto-encoders for co-expression prediction is a methodological advance, but it does not change the fundamental limitation that co-expression is an indirect evidence type. When you see a co-expression edge in your network, consider whether the two proteins are expected to be co-regulated based on your experimental system.

### Low Correlation between Transcriptomics and Proteomics

The [lens-induced myopia study](https://pubmed.ncbi.nlm.nih.gov/37819745) found low correlation between transcriptomic and proteomic data. This is a common observation in multi-omics experiments and reflects the many regulatory steps between mRNA and protein abundance. When you integrate STRING analysis across data types, do not expect the same proteins to be significant in both data sets.

The practical implication is that pathway-level analysis is often more informative than protein-level comparison. The myopia study found that the transcriptomic and proteomic data complemented one another in KEGG pathway annotation, with metabolic and human disease pathways correlated with the myopia-forming process. This pathway-level convergence is more meaningful than individual protein overlap.

## Quality Controls and Records for STRING Analysis

### Documentation Requirements

Maintain a record of every STRING analysis you perform. The record should include the STRING version, the date, the organism, the input identifier type, the number of input proteins, the number of mapped proteins, the confidence threshold, and the network type. This information is essential for reproducibility and for responding to reviewer requests.

The [nf-core documentation](https://nf-co.re/docs) emphasizes community pipeline standards, usage, configuration, and reproducible workflow context. Applying similar standards to your STRING analysis ensures that your results can be reproduced by other researchers.

### Validation of Key Findings

Before building a publication figure around a STRING network, validate the key interactions. Check whether the interactions are supported by experimental evidence in the STRING database and whether the confidence scores are high. If a critical edge in your network has low confidence, consider whether you should highlight it as a prediction instead of an established interaction.

For proteomics data, validation can also include checking whether the key proteins were identified with sufficient spectral evidence. A protein identified by a single peptide is less reliable than a protein identified by multiple unique peptides, and this reliability should be considered when interpreting its role in the network.

### Version Control for Reproducibility

STRING updates its database regularly, and the same query can produce different results in different versions. If you are writing a manuscript or thesis, record the exact STRING version you used. If you repeat the analysis after a STRING update, document the version change and check whether the conclusions are affected.

The [Carpentries lessons](https://carpentries.org/lessons) provide foundational training in data management and version control that is applicable to bioinformatics analyses. Using version control for your analysis scripts and recording the STRING version in your methods section are basic reproducibility practices.

## Professional Escalation Criteria for STRING Results

### When to Seek Bioinformatics Support

If your STRING analysis produces unexpected results that you cannot explain, seek help from a bioinformatics specialist. Specific situations that warrant escalation include: a very low mapping rate for your input identifiers, enrichment results that are dominated by generic terms such as "protein binding," or a network that is so dense that no meaningful clusters can be identified.

The [EMBL-EBI Training](https://www.ebi.ac.uk/training) program provides bioinformatics learning pathways and practical analysis education that can help you build the skills to troubleshoot common problems. If you are repeatedly encountering the same issues, investing time in structured training may be more efficient than seeking help for each individual analysis.

### When to Question the Experimental Data

Sometimes the problem is not in the STRING analysis but in the input data. If your differentially abundant protein list contains many proteins that are not biologically plausible for your experimental system, review the mass spectrometry data quality. Check the number of peptides per protein, the coefficient of variation between replicates, and the presence of contaminants.

The [Bioconductor project](https://bioconductor.org/) provides official package, workflow, installation, and reproducible genomic-analysis documentation. Many proteomics quality-control packages are available through Bioconductor, and these can help you assess whether your input data meet the standards for meaningful network analysis.

### When to Consult a Domain Expert

STRING analysis identifies pathways and interactions, but interpreting the biological significance requires domain expertise. If your network implicates a pathway that is outside your area of expertise, consult a researcher who studies that pathway. For example, if your proteomics data from a cancer study implicate complement and coagulation proteins, a hematology or immunology expert can help you interpret the finding.

The [thymoma study](https://pubmed.ncbi.nlm.nih.gov/36991043) identified complement and coagulation cascade proteins as key activators, including von Willebrand factor, coagulation factor V, and vitamin K-dependent protein C. Interpreting this finding in the context of thymoma biology required knowledge of both tumor immunology and coagulation biology. A domain expert can help you determine whether the finding is mechanistically plausible or an artifact of serum protein abundance.

## Practical Workflow for STRING Analysis of Proteomics Data

### Step 1: Prepare the Protein List

Export the list of differentially abundant proteins from your proteomics software. Include the gene symbol or UniProt identifier, the log2 fold change, and the adjusted p-value. Remove contaminants and reverse hits. Record the total number of quantified proteins for use as the background set.

### Step 2: Map Identifiers

Submit your protein list to STRING and check the mapping statistics. If any identifiers fail to map, investigate the cause. Convert identifiers if necessary using the STRING mapping tool or an external resource such as [NCBI](https://www.ncbi.nlm.nih.gov/).

### Step 3: Select Organism and Confidence Threshold

Confirm that the correct organism is selected. Choose a confidence threshold based on your research question. For exploratory analysis, medium confidence is a reasonable starting point. For confirmatory analysis, high confidence is more appropriate.

### Step 4: Build the Network

Generate the network and inspect the edge colors to understand the evidence channels. Identify clusters and examine the proteins within each cluster. Note any hub proteins and consider whether their centrality reflects biological importance or literature bias.

### Step 5: Run Enrichment Analysis

Run enrichment analysis against Gene Ontology, KEGG, and other categories. Use the appropriate background set. Record the false discovery rate values and select significant terms using a predefined threshold.

### Step 6: Export and Document

Export the network image and the tabular edge data. Record the STRING version, the date, the input parameters, and the mapping statistics. Save all files in a project directory with clear naming conventions.

### Step 7: Validate and Interpret

Validate the key interactions by checking the experimental evidence in STRING. Interpret the enriched pathways in the context of your experimental system. Consult domain experts if the findings implicate pathways outside your expertise.

## Records and Measurements for STRING Analysis

### What to Record

For each STRING analysis, record the following information in your laboratory notebook or electronic lab notebook:

- STRING version and access date
- Organism and genome annotation version
- Input identifier type and number of proteins
- Number of mapped and unmapped proteins
- Confidence threshold and network type
- Background set used for enrichment
- Enrichment results with false discovery rate values
- Network clustering parameters
- Export file names and formats

### What to Measure

The key measurements for evaluating a STRING analysis are the mapping rate, the number of significant enrichment terms, and the network density. A mapping rate below 90 percent warrants investigation. A network with very few edges at high confidence may indicate that your proteins are not well characterized. A network with too many edges may be difficult to interpret.

The number of significant enrichment terms depends on the size of your protein list and the background set. A list of 50 proteins typically produces fewer significant terms than a list of 500 proteins. Compare your results with published studies that used similar experimental designs to calibrate your expectations.

## Safety and Ethical Context for STRING Analysis

### Data Sharing and Privacy

Proteomics data from human samples may contain sensitive information. When you submit protein lists to STRING, you are sending data to an external server. For human data, ensure that your institutional review board approval covers the analysis and that you are not sharing identifiable information.

The [UK Biobank study](https://pubmed.ncbi.nlm.nih.gov/38710337) of plasma proteomic biomarkers involved 28,974 participants and used OLINK proteomics to measure 1,463 plasma proteins. Studies of this scale require careful attention to data governance and participant privacy. Your STRING analysis should follow the same principles, even at a smaller scale.

### Reproducibility as an Ethical Obligation

Publishing a STRING analysis without documenting the parameters is a reproducibility failure. Other researchers cannot reproduce your network or enrichment results if they do not know the STRING version, the confidence threshold, and the background set. The [nf-core documentation](https://nf-co.re/docs) and the [Galaxy Training Network](https://training.galaxyproject.org/) both emphasize reproducibility as a core principle of bioinformatics analysis.

### Responsible Interpretation

The interpretation of STRING results has ethical dimensions. Overinterpreting a network as proof of a mechanism can mislead other researchers and waste resources. The low correlation between transcriptomics and proteomics observed in the [myopia study](https://pubmed.ncbi.nlm.nih.gov/37819745) is a reminder that each data type provides a partial view of biology. Your STRING analysis should be presented as one line of evidence among several.

## Frequently Asked Questions

### What is the difference between a physical interaction and a functional association in STRING?

A physical interaction means that two proteins directly bind to each other, as demonstrated by experiments such as co-immunoprecipitation or yeast two-hybrid. A functional association means that two proteins participate in the same biological process or pathway without necessarily binding directly. STRING integrates both types of evidence, and the full network includes all interaction types while the physical subnetwork includes only direct binding. For proteomics data, the full network is usually more informative because it captures pathway relationships that are not based on direct binding.

### How do I choose the right confidence threshold for my STRING network?

The confidence threshold depends on your research question. For exploratory analysis where you want to generate hypotheses, medium confidence is a reasonable starting point because it balances sensitivity and specificity. For confirmatory analysis where you want to highlight well-established interactions, high confidence produces a sparser network with fewer false-positive edges. Document the threshold you use, because it affects the network topology and the interpretation of clusters.

### What should I do if many of my protein identifiers do not map in STRING?

First, check whether you are using the correct identifier format. STRING accepts gene symbols, UniProt identifiers, Ensembl identifiers, and several other formats. If the format is correct, check whether the identifiers are current. Protein databases are updated regularly, and obsolete accession numbers may fail to map. Remove contaminant proteins from your list before submitting, because many search engines append a contaminant prefix to common laboratory contaminants. If more than 10 percent of your list fails to map, investigate the cause before proceeding with the analysis.

### What background set should I use for enrichment analysis in STRING?

The background set should match the scope of your experiment. If your mass spectrometry experiment quantified the whole proteome, the entire genome is an appropriate background. If your experiment only detected a subset of the proteome, use the detected protein set as the background. For targeted approaches such as OLINK, use the panel of measured proteins. Using the wrong background set can bias the enrichment results toward the measurable proteins and produce false-positive pathway enrichment.

### Why do my transcriptomics and proteomics data show low correlation?

Low correlation between mRNA and protein abundance is common and reflects the many regulatory steps between transcription and protein abundance, including translation efficiency, protein degradation, and post-translational modifications. The [lens-induced myopia study](https://pubmed.ncbi.nlm.nih.gov/37819745) found low correlation between transcriptomic and proteomic data, but the pathway-level analysis revealed that the two data sets complemented one another. Instead of comparing individual proteins, compare the enriched pathways and interaction networks from each data set.

### Can I use STRING for organisms that are not in the database?

STRING version 12.0 allows you to create, browse, and analyze a full interaction network for any novel genome of interest by submitting its complement of encoded proteins. The interaction predictions are transferred from well-studied organisms using hierarchical orthology information. This feature enables network analysis for any sequenced genome, but the transferred interactions are less reliable than direct experimental evidence from the same organism.

### How should I export STRING results for publication?

Export the network image in a vector format such as SVG for publication, because vector formats scale without losing resolution. Include a legend that explains the edge colors and the confidence score. Also export the tabular edge data for supplementary materials and for downstream analysis. Record the STRING version, the organism, the input identifiers, the confidence threshold, and the date of analysis so that other researchers can reproduce your results.

### What are the main limitations of STRING for proteomics data?

STRING coverage depends on the available interaction data for each organism, and well-studied organisms have more extensive records than less-studied organisms. The database is biased toward well-studied proteins, which can appear as hubs in the network even when their biological importance in your specific experiment is unclear. Co-expression predictions from proteomics data are indirect evidence and do not prove physical interaction. The low correlation between transcriptomics and proteomics means that multi-omics comparisons should focus on pathway-level analysis instead of individual protein overlap.

## Related Bioinformatics Guides

- [STRING Database and Protein-Protein Interaction Networks](/knowledge/bioinformatics/string-database-and-protein-protein-interaction-networks)
- [Genomic Data Analysis Tools: A Comparative Guide for Researchers](/knowledge/bioinformatics/genomic-data-analysis-tools-a-comparative-guide-for-researchers)
- [How to Interpret Gene Set Enrichment Analysis Results](/knowledge/bioinformatics/how-to-interpret-gene-set-enrichment-analysis-results)
- [Proteomics Mass Spectrometry: From Sample Preparation to Data Analysis](/knowledge/bioinformatics/proteomics-mass-spectrometry-from-sample-preparation-to-data-analysis)
- [TMT Proteomics: Experimental Design, Labeling, and Data Analysis](/knowledge/bioinformatics/tmt-proteomics-experimental-design-labeling-and-data-analysis)

## References and Further Reading

- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Bioconductor](https://bioconductor.org/). Bioconductor Project.
- [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.
- [nf-core Documentation](https://nf-co.re/docs). nf-core.
- [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.
- [The STRING database in 2023: protein-protein association networks and functional enrichment analyses for any sequenced genome of interest.](https://pubmed.ncbi.nlm.nih.gov/36370105). Nucleic acids research, 2023.
- [Plasma proteomic biomarkers and the association between poor cardiovascular health and incident dementia: The UK Biobank study.](https://pubmed.ncbi.nlm.nih.gov/38710337). Brain, behavior, and immunity, 2024.
- [Integrative Transcriptome and Proteome Analyses Elucidate the Mechanism of Lens-Induced Myopia in Mice.](https://pubmed.ncbi.nlm.nih.gov/37819745). Investigative ophthalmology & visual science, 2023.
- [ProteomicsDB.](https://pubmed.ncbi.nlm.nih.gov/29106664). Nucleic acids research, 2018.
- [Proteomics analysis of serum from thymoma patients.](https://pubmed.ncbi.nlm.nih.gov/36991043). Scientific reports, 2023.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.