Human Protein Atlas
The Human Protein Atlas (HPA) is a publicly accessible, antibody‑based resource that maps the expression and spatial distribution of proteins across human tissues, cell lines, and organs. It integrates transcriptomics (RNA‑seq) data with immunohistochemistry images to provide a multi‑layered view of the proteome. This guide is intended for life science researchers, bioinformatics analysts, and students who need to retrieve, interpret, and quality‑check protein localization data. If you are planning experiments involving tissue‑specific expression, biomarker discovery, or validation of antibody specificity, the HPA is a key reference tool. NCBI Bookshelf contains background chapters on proteomics that help contextualize the atlas, while EMBL‑EBI Training offers modular tutorials on navigating the resource.
The HPA covers nearly 90% of all human protein‑coding genes and provides expression data at the tissue, cell, and subcellular levels. Its six main sections (Tissue, Cell, Pathology, Brain, Blood, and Metabolic) allow users to search by gene or protein. Understanding how to filter by antibody validation score, tissue type, or pathological condition is critical for drawing reliable biological conclusions. Galaxy Training Network offers workflows that combine HPA downloads with transcriptomic analyses, and Bioconductor packages such as hpaVis enable programmatic access for large‑scale comparisons.
At a Glance
| Feature | Description |
|---|---|
| Data types | Immunohistochemistry (IHC) images, immunofluorescence, RNA‑seq expression (TPM), protein‑protein interaction networks |
| Primary methods | Antibody‑based protein detection, RNA‑seq from 37 tissue types (FANTOM5 / GTEx) |
| Spatial resolution | Tissue‑level (IHC on tissue microarrays), subcellular (immunofluorescence on cell lines) |
| Validation tiers | Approved, Supported, Uncertain , based on antibody specificity and reproducibility |
| Access | Free web portal (proteinatlas.org) and downloadable files |
| Primary use cases | Exploring tissue‑specific expression, identifying candidate biomarkers, comparing protein and mRNA levels |
Decision Points and Criteria
Before using the HPA, decide which data layer matches your question. The atlas contains both antibody‑based protein data (IHC scoring) and transcriptomics (RNA‑seq). These two sources often correlate but can diverge due to post‑transcriptional regulation. For hypothesis generation, RNA‑seq values give a quantitative baseline, while IHC images provide spatial context.
When to rely on antibody‑based data
- You need to know where a protein resides within a tissue (e.g., cytoplasm vs. nucleus).
- You are validating an antibody for immunohistochemistry or immunofluorescence.
- You work on proteins with low mRNA abundance that may still be detectable by sensitive antibodies.
When to rely on transcriptomics
- You want a quantitative measure across many tissues for a first pass.
- You plan to compare your own RNA‑seq results with a reference.
- You are studying genes that lack validated antibodies.
When to combine both
- You want to assess whether mRNA levels predict protein presence. Large discrepancies may indicate post‑translational regulation or antibody cross‑reactivity.
- You are building a multi‑omics model and need orthogonal evidence.
The HPA assigns a reliability score to each antibody: Enhanced, Supported, Approved, or Uncertain. Filter for “Approved” or “Supported” before drawing strong conclusions. NCBI Sequence Read Archive can provide raw RNA‑seq data that complements HPA’s processed expression values.
Practical Workflow
Follow this sequence to extract and interpret HPA data without common pitfalls.
Step 1: Identify your gene of interest
Use the search bar on the HPA homepage with official gene symbol (e.g., EGFR) or UniProt ID. Review the gene summary page, which links to all six sections.
Step 2: Check antibody validation
Click the “Antibodies” tab to see the antibody ID, reliability score, and validation against knockout cells and protein arrays. Only use data from antibodies with a score of “Supported” or higher.
Step 3: Examine tissue expression
Go to “Tissue Atlas.” The heatmap shows RNA‑seq expression (transcripts per million, TPM) alongside IHC scores (not detected, low, medium, high). Click a tissue to view the corresponding IHC image. Download the high‑resolution image for qualitative assessment.
Step 4: Compare subcellular location
Switch to “Cell Atlas” to see immunofluorescence images for three standard cell lines (U‑2 OS, A‑431, U‑251 MG). The subcellular location is annotated as nuclear, cytoplasmic, membranous, or organelle‑specific. Note that a single cell line may not represent the in vivo context.
Step 5: Explore pathology associations
The “Pathology Atlas” links protein expression to survival data for 17 cancer types. Use this cautiously: associations do not prove causality, and the sample size per tumor type varies. A roadmap to generate renewable protein binders to the human proteome emphasizes that antibody specificity remains the largest source of variation in such studies.
Step 6: Download data for your analysis
Go to “Downloads” to retrieve ready‑to‑use files: RNA‑seq TPM matrix, IHC scores, or subcellular location tables. For programmatic access, use the R package hpaVis from Bioconductor. Verify file version and release date.
Step 7: Perform quality checks
- Compare IHC scores with external databases such as GTEx or The Human Proteome Map.
- Check for multiple antibodies targeting the same protein. Discordant results often indicate off‑target binding.
- Ensure your tissue of interest is represented. The atlas covers 44 normal tissues and 20 cancer types, but some rare tissues may be absent.
Common Mistakes
- Using unvalidated antibodies for high‑stakes conclusions. Always filter by antibody reliability. A “Uncertain” antibody may still bind the correct target but often shows cross‑reactivity.
- Ignoring the semi‑quantitative nature of IHC scores. The scoring (low, medium, high) is categorical and based on pathologist annotation. It is not equivalent to absolute protein concentration. NCBI Bookshelf details how IHC quantification is inherently less precise than mass spectrometry.
- Equating RNA and protein expression without caution. Abundant mRNA does not guarantee abundant protein due to degradation, translation efficiency, or secretion. Conversely, a high IHC score with low RNA may indicate an antibody artifact.
- Overinterpreting subcellular location from a single cell line. The Cell Atlas uses a few model lines, location may differ in primary cells or disease states. Use the data as a hypothesis, not a fact.
- Neglecting batch effects in pathology data. The Pathology Atlas assembles tissues from multiple cohorts. Check the number of samples per cancer type before making survival correlations.
Limits of Interpretation
The HPA is an indispensable resource, but it has inherent constraints that users must acknowledge.
Antibody specificity and bias. The entire atlas depends on the quality of commercial antibodies. Even “Approved” antibodies can bind off‑target proteins in different tissues or fixatives. The roadmap paper A roadmap to generate renewable protein binders to the human proteome notes that only about 50% of antibodies in the HPA have passed the highest validation criteria. Therefore, any finding based on a single antibody should be replicated with an orthogonal method (e.g., Western blot, mass spectrometry).
Tissue representation and sampling. The atlas includes 44 normal tissues, but each tissue is represented by a single tissue microarray donor. This ignores inter‑individual variation due to age, sex, or genetic background. For some tissues (e.g., adrenal gland) only one block is available. The EMBL‑EBI Training resource warns users against generalizing expression levels to all individuals.
Semi‑quantitative nature of protein data. IHC scores are ordinal, not continuous. A “high” score in one tissue versus “low” in another does not reveal fold‑change. RNA‑seq TPM values are quantitative but come from a single RNA‑seq experiment and may not reflect the same donor as the IHC. Comparability across data types is limited.
Subcellular localization in cell lines versus tissues. The Cell Atlas uses cultured cells that often differ from the in vivo state. For example, proteins normally secreted may accumulate in the Golgi in cell culture. Always validate live‑cell or primary tissue localization with independent methods.
Pathology associations are correlative. The survival data in the Pathology Atlas are unadjusted for confounders such as tumor stage or treatment. A high expression of a protein may correlate with better survival without being a direct protector. The Galaxy Training Network offers workflows to adjust survival curves with clinical covariates, but HPA’s online interface does not.
Frequently Asked Questions
1. How can I download the entire Human Protein Atlas dataset?
You can download bulk data from the “Downloads” page on the HPA website. Files are available as tab‑separated values for RNA‑seq expression, IHC scores, and subcellular location. For RNA‑seq data, the matrix contains TPM values for each tissue. Updated releases occur approximately once per year. Check the version number before publishing.
2. What does “Enhanced” antibody reliability mean?
An “Enhanced” antibody has passed the European Bioinformatics Institute’s validation pipeline using knockout cells, protein arrays, and consistency with RNA data. This is the highest confidence level. Use “Supported” antibodies as a secondary choice. Antibodies marked “Uncertain” lack sufficient validation and should not be used without independent confirmation.
3. Why does the same gene show conflicting IHC scores among different tissues?
Tissue‑specific post‑translational modifications, alternative splicing, or differences in epitope accessibility can cause variable antibody binding. Additionally, some tissues may have inherently higher protein expression. The Bioconductor package hpaVis can help you plot correlations and flag inconsistencies.
4. Can I use HPA data to determine whether a protein is secreted?
The HPA does not directly annotate secretion pathways, but you can infer extracellular location from the subcellular classification “secreted” in the Cell Atlas or from IHC staining in glandular tissues. A better approach is to complement HPA with signal peptide predictions (e.g., SignalP) and data from the Human Secretome Atlas. The NCBI Sequence Read Archive can provide RNA‑seq from exosomal fractions if you need deeper evidence.
References and Further Reading
- Human Protein Atlas official portal. https://www.proteinatlas.org/
- Uhlén M, Fagerberg L, Hallström BM, et al. Tissue‑based map of the human proteome. Science. 2015,347(6220):1260419. https://pubmed.ncbi.nlm.nih.gov/25613900/
- Thul PJ, Åkesson L, Wiking M, et al. A subcellular map of the human proteome. Science. 2017,356(6340):eaal3321. https://pubmed.ncbi.nlm.nih.gov/28495876/
- Colwill K, Graslund S, Tyers M, et al. A roadmap to generate renewable protein binders to the human proteome. Nat Methods. 2011,8(7):551‑558. https://pubmed.ncbi.nlm.nih.gov/21572409/
- NCBI Bookshelf , free textbooks covering proteomics methods and antibody validation.
- EMBL‑EBI Training , online courses on the Human Protein Atlas and related bioinformatics.
- Galaxy Training Network , workflows for integrating HPA data with RNA‑seq analysis.
- Bioconductor , software packages for programmatic access and visualization of HPA data.
- NCBI Sequence Read Archive , repository for raw sequencing data used in HPA transcriptome generation.
- Digre A, Lindskog C. The Human Protein Atlas , spatial localization of the human proteome in health and disease. Protein Sci. 2021,30(1):218‑233. https://pubmed.ncbi.nlm.nih.gov/33350026/