Zubair Khalid

Virologist/Molecular Biologist | Veterinarian | Bioinformatician

Conventional & Molecular Virology • Vaccine Development • Computational Biology

Dr. Zubair Khalid is a veterinarian and virologist specializing in conventional and molecular virology, vaccine development, and computational biology. Dedicated to advancing animal health through innovative research and multi-omics approaches.

Dr. Zubair Khalid - Veterinarian, Virologist, and Vaccine Development Researcher specializing in Computational Biology, Multi-omics, Animal Health, and Infectious Disease Research

Category: Guides

Protein Atlas

A protein atlas is a comprehensive, systematically curated resource that maps the expression, localization, and abundance of proteins across different tissues, cell types, developmental stages, and disease states. For biologists, bioinformaticians, and clinicians who need to understand where and when a protein is present, the protein atlas serves as a reference to generate hypotheses, validate findings, and contextualize omics data. This guide provides a source bounded framework for using protein atlas resources, with practical steps to extract reliable information while avoiding common pitfalls. Begin by visiting the NCBI Bookshelf for foundational reviews on proteome mapping.

Protein atlases integrate data from multiple methodologies, including antibody based immunohistochemistry, mass spectrometry, and transcriptomics. The best known example, the Human Protein Atlas (HPA), covers more than 90% of all human protein coding genes across 44 normal tissues, 47 cell lines, and 20 cancer types. Whether you study development, disease biomarkers, or drug targets, a protein atlas tells you if a protein of interest is expressed in your system of choice and at what relative level. For training on how to explore these resources, consult the EMBL-EBI Training portal which offers modules on proteomics data interpretation.

At a Glance

Aspect Description
What a protein atlas provides Tissue and cell specific protein expression maps with subcellular localization
Main data types Immunohistochemistry images, RNA expression, mass spectrometry based proteomics
Primary user Molecular biologists, disease researchers, drug developers
Key resource The Human Protein Atlas (proteinatlas.org)
Validation level Each antibody is validated against multiple criteria including protein array and western blot
Access Free online browsable database with downloadable data files
Update frequency Quarterly releases with continuous curation

Core Concepts and Decision Points

A protein atlas is not a single homogeneous database. Different tissue atlases, pathology atlases, and cell atlases exist for specific research questions. Understanding the core concepts helps you choose the correct resource.

Tissue atlas versus cell atlas. A tissue atlas shows protein distribution in whole organ sections, revealing cell populations and their relative staining intensity. A cell atlas, like the single cell type atlas in HPA, focuses on individual cells and can distinguish rare cell types. If you need to know whether your protein is found in hepatocytes or in Kupffer cells, a cell atlas is more informative. For an overview of such cellular resolution studies, see the Galaxy Training Network which provides workflows for single cell proteomics analysis.

Transcript versus protein correlation. RNA levels are often used as a proxy for protein abundance, but the correlation is moderate, especially in tissues with high post transcriptional regulation. A decision point arises: when RNA data from public repositories like the NCBI Sequence Read Archive suggests expression, always confirm with protein atlas immunohistochemistry or mass spectrometry data. The HPA explicitly provides both RNA consensus and protein evidence so you can compare directly.

Choice of species. Most comprehensive atlases are for human and mouse. For other organisms, such as wheat or zebrafish, partial atlases exist. A recent study on durum wheat genomics Durum Wheat cv. Svevo Reference Genome Rel.2.0 illustrates how protein atlases are being extended to crops. When working with non model organisms, consider referencing such species specific resources or using cross species antibodies with caution.

Decision criteria for using a protein atlas:

  • Are you asking where a protein is expressed? Use a tissue atlas.
  • Do you need subcellular location? Use the subcellular section of HPA.
  • Is your protein a secreted factor? Check the secretome predicted by signal peptide tools.
  • Are you comparing cancer versus normal? Use the pathology atlas.

Practical Workflow or Implementation Sequence

Follow these steps to query a protein atlas and extract reliable answers.

  1. Identify your protein of interest. Use the official gene symbol (e.g., TP53, EGFR). Avoid aliases until you have confirmed the entry. Go to the Human Protein Atlas and search the gene name.

  2. Examine the summary tab. This shows the main expression summary: “Tissue expression” with a bar chart of normalized expression (nTPM) across organs. Note that RNA expression is shown first, then protein expression via immunohistochemistry.

  3. Review the immunohistochemistry images. Each tissue has a stained slide with multiple annotations. Look for the staining intensity (negative, weak, moderate, strong) and quantity (less than 25%, 25 75%, more than 75% of cells). The combination tells you both prevalence and per cell abundance.

  4. Check the subcellular section. Click on the cell line or tissue data to see where the protein localizes. This is critical because a protein may be nuclear in one cell type and cytoplasmic in another. A recent atlas of RNA polymerase III activity An RNA polymerase III tissue and tumor atlas demonstrates context specific localization changes.

  5. Validate with other data sources. Cross reference with the Bioconductor package HPAanalyze which allows programmatic access to HPA data. This step is especially important if you plan to use the protein atlas to guide antibody selection or design experiments. Run quality checks as described below.

  6. Download the data. For large scale analysis, download the complete dataset (normal tissue data, pathology data, subcellular location data) from the HPA download page. Files are in tab separated format and can be imported into R or Python.

  7. Document your findings. Record the antibody ID used, the tissue batch, and the staining pattern. This documentation ensures reproducibility and helps future comparisons.

Quality Checks

Every protein atlas data point carries inherent uncertainty. Apply these quality checks before using the information.

  • Antibody validation summary. The HPA grades each antibody as “supported” or “approved” based on multiple validation methods. Look for the orange or green badge. If a protein is based on a single antibody with weak validation, treat the result as preliminary.

  • Consistency between antibodies. Some genes have two or more antibodies targeting different epitopes. Compare the staining patterns. If they differ, investigate the discrepancy by checking binding specificity profiles.

  • RNA protein concordance. In a given tissue, if RNA expression is high but protein staining is absent, consider post transcriptional regulation or technical issues such as epitope masking. Conversely, if protein is detected without RNA, check for alternative splicing or antibody cross reactivity.

  • Replicate reproducibility. The HPA shows multiple tissue samples for the same organ (e.g., three different donors). If staining varies markedly between donors, the protein may be expressed in a subset of individuals or under specific conditions. The pan disease analysis of rheumatic autoimmune diseases Pan disease blood protein profiles illustrates how population level variability affects atlas data.

Common Mistakes

Even experienced researchers make these errors when using protein atlases.

Mistake 1: Assuming absence means no expression. A negative immunohistochemistry result does not prove the protein is absent. The antibody may fail to detect the protein in that tissue due to fixation, low abundance, or splicing variants. Use multiple detection methods before concluding absence.

Mistake 2: Overinterpreting intensity as absolute concentration. Staining intensity (weak, moderate, strong) is ordinal, not quantitative. It compares relative levels within the same slide batch but not across different antibodies. Do not use intensity values for statistical modeling without normalization.

Mistake 3: Ignoring subcellular compartment. A protein may be functional in the nucleus but reported in the cytoplasm for certain cell lines. Always check the subcellular location for the cell type relevant to your question. A study on FABP2 in renal cell carcinoma Radiologic and lipid metabolism imaging features highlights how subcellular mislocalization can change interpretation.

Mistake 4: Citing an atlas without context. When referencing a protein atlas finding, include the antibody ID, tissue donor, and data version. General statements like “expressed in liver” are insufficient for rigor. Provide the link to the specific entry.

Mistake 5: Skipping the quality section. Many users jump directly to the images without reviewing the antibody validation summary. This can lead to using unreliable data.

Limits and Uncertainty

Protein atlases have well recognized limitations that should temper any strong conclusion.

  • Sample bias. Most atlases use tissue from a limited number of donors (often one to three per tissue). This does not capture the full inter individual variation. A cross ancestry atlas of infectious disease A cross ancestry genetic atlas demonstrates that protein expression can differ across populations, yet most atlases are derived from individuals of European descent.

  • Detection threshold. Immunohistochemistry can miss low abundance proteins, especially transcription factors and signaling molecules. Mass spectrometry based proteomics has a wider dynamic range but still struggles with low copy number proteins.

  • Epitope availability. Tissue fixation (formalin fixed paraffin embedded) can mask epitopes, leading to false negatives. The HPA uses a standardized protocol, but some antigens remain inaccessible. Check the “enhanced” validation category if available.

  • Lack of quantitative precision. Protein atlas data are semi quantitative. They are excellent for discovering presence or absence and relative differences within a tissue, but not for fold change comparisons across tissues.

  • Cell type resolution in bulk tissues. Even careful manual annotation of immunohistochemistry slides cannot fully resolve protein expression in all cell types, especially when cells are intermixed. Single cell proteomic technologies are emerging but not yet standard in atlases. A recent immunopeptide study Identification of immunopeptides in chondrosarcoma used a combination of proteomics and transcriptomics to overcome this limitation.

Frequently Asked Questions

Q: Can I trust the RNA expression data from a protein atlas when the protein data disagree? RNA and protein levels often diverge due to post transcriptional regulation. Trust protein data if supported by validated antibodies. Use RNA data as a complementary measure, but prioritize direct protein evidence when possible.

Q: How often is the Human Protein Atlas updated? The HPA releases a new version quarterly. Major updates occur yearly with new tissues, antibodies, and methods. Always note the version you used, as older versions may be outdated.

Q: Is the protein atlas useful for non human species? The HPA primarily covers human and mouse. For other species, check dedicated resources: for example, the durum wheat reference genome mentioned earlier provides a foundation for building a wheat protein atlas. Most model organisms have tissue specific proteome databases, but they are less comprehensive.

Q: What should I do if my protein of interest is not in the atlas? It may be uncharacterized or not yet targeted by validated antibodies. You can request data via the HPA antibody submission program, or search other resources such as UniProt and the NCBI Sequence Read Archive for evidence at the RNA level.

References and Further Reading

Related Articles