Zubair Khalid

Virologist/Molecular Biologist | Veterinarian | Bioinformatician

Conventional & Molecular Virology • Vaccine Development • Computational Biology

Dr. Zubair Khalid is a veterinarian and virologist specializing in conventional and molecular virology, vaccine development, and computational biology. Dedicated to advancing animal health through innovative research and multi-omics approaches.

Dr. Zubair Khalid - Veterinarian, Virologist, and Vaccine Development Researcher specializing in Computational Biology, Multi-omics, Animal Health, and Infectious Disease Research

Section: Infrastructure, Cloud & Policy

Genomic Data Repositories: Navigating Public Databases for Research

Public genomic repositories now hold sequence, expression, and variation data from thousands of studies across human health, agriculture, microbiology, and ecology. Researchers who need to retrieve, compare, or reanalyze genomic data face a practical challenge: each repository has distinct data types, access policies, submission formats, and retrieval tools. This article provides a working orientation to major public genomic data repositories, with emphasis on the National Center for Biotechnology Information (NCBI) resources, the European Bioinformatics Institute (EBI) resources, and the National Cancer Institute Genomic Data Commons (GDC). The focus is on practical decisions: where to look for a given data type, how to interpret access requirements, and how to avoid common retrieval and reuse errors.

At a Glance

The table below summarizes the major repositories covered in this article, the data types they hold, and their access characteristics. Use this table as a first filter when deciding where to search for a specific dataset.

Repository Primary Data Types Access Model Typical Use Cases
NCBI Sequence Read Archive (SRA) Raw sequencing reads from DNA and RNA sequencing Open access for most datasets, controlled access for human data under specific policies Variant calling, metagenomics, transcriptomics, genome assembly
NCBI Gene Expression Omnibus (GEO) Functional genomics data: microarray and high-throughput sequencing expression data Open access, data are freely downloadable Differential expression analysis, meta-analysis of expression studies
ENA (European Nucleotide Archive) Raw sequencing reads, assembled genomes, and sequence annotations Open access for most data, controlled access where required by data policies Cross-continental data retrieval, sequence alignment, genome browsing
NCI Genomic Data Commons (GDC) Cancer genomics data including DNA sequencing, RNA sequencing, methylation, and clinical metadata Open access for many datasets, controlled access for some human data Cancer genomics comparisons, harmonized analysis across cohorts
IPD-IMGT/HLA Database Human leukocyte antigen (HLA) allele sequences Open access Transplantation matching, immunogenetics, population studies
iProX Mass spectrometry-based proteomics data Open access after release, full member of ProteomeXchange consortium Proteomics data sharing and retrieval, particularly for China-based studies
GeOMe Field and sampling event metadata associated with genetic samples Open access Ecological genomics, specimen tracking, environmental context

Understanding the Repository Landscape

Genomic data repositories differ in scope, governance, and technical architecture. The NCBI hosts a family of integrated databases including the Sequence Read Archive, Gene Expression Omnibus, and numerous variation and reference resources. The EBI provides complementary services including training materials and the European Nucleotide Archive. The GDC focuses specifically on cancer genomics and applies uniform processing pipelines to harmonize data across projects.

The FAIR Guiding Principles provide a framework for evaluating repository quality. These principles state that data should be Findable, Accessible, Interoperable, and Reusable. When you select a repository for data retrieval or deposition, assess whether the repository supports persistent identifiers, clear access conditions, standardized metadata, and explicit reuse licenses. The FAIR principles were published in Scientific Data and have become a common benchmark for repository evaluation.

Repository Governance and Data Policies

Genomic data repositories operate under governance frameworks that balance open science with participant privacy. The NIH Genomic Data Sharing Policy establishes expectations for data deposition and access for NIH-funded research. This policy requires that large-scale genomic data be shared through designated repositories and that access to certain human data be controlled through authorized access mechanisms.

Commercial interests in genomic data introduce additional governance considerations. A systematic literature review of public attitudes toward biobank and genomic data repositories found that commercial involvement raises concerns about consent, privacy, data security, trust in science, benefit sharing, and ownership and control of health data. These findings indicate that repository governance and access policies must be carefully designed to maintain public trust. When you use data from a repository, review the access conditions and data use restrictions that apply to each dataset.

Core Principles for Repository Use

Effective use of genomic data repositories depends on understanding several core principles that govern how data are organized, accessed, and interpreted.

Data Types and Their Organizational Units

Each repository organizes data according to its own logical model. The Gene Expression Omnibus uses three central entities: platforms, samples, and series. A platform is a list of probes that define what set of molecules may be detected. A sample describes the set of molecules being probed and references a single platform used to generate its molecular abundance data. A series organizes samples into meaningful datasets that make up an experiment. Understanding these entities helps you navigate GEO records and retrieve the specific data you need.

The Sequence Read Archive organizes data by study, sample, experiment, and run. A study corresponds to a research project, a sample represents the biological material, an experiment describes the sequencing protocol, and a run contains the actual sequence data. When you retrieve data from SRA, you typically start with a study accession and then select the runs relevant to your analysis.

Metadata Quality and Its Consequences

Metadata quality varies across repositories and submissions. Some datasets include detailed experimental protocols, sample characteristics, and processing information. Others contain minimal metadata that limits their reuse value. The Genomic Observatories Metadatabase was developed specifically to address the need for field and sampling event metadata associated with genetic samples, recognizing that ecological and environmental context is essential for interpreting genomic data.

When you evaluate a dataset for reuse, examine the metadata thoroughly before downloading large files. Check whether the metadata describes the organism, tissue type, experimental conditions, sequencing platform, and data processing steps. If critical metadata are missing, contact the data submitter or consider whether the dataset is suitable for your purpose.

Data Harmonization and Processing Consistency

Data generated by different laboratories and processed with different pipelines may not be directly comparable. The GDC addresses this challenge by applying uniform processing pipelines to cancer genomics data. Harmonized data from the GDC can serve as controls or comparisons in compendium analyses with user data, avoiding the expense and time of generating additional datasets. However, for these comparisons to be useful, the user must process their new data in the same manner as the repository data.

The GDC pipelines that describe the entire analysis workflow were originally published as text-based standard operating procedures. These have been converted into downloadable and executable formats, including containerized and interactive graphical workflows. These workflows can be applied to reproducibly process user data and to harmonize datasets across repositories. When you plan a comparative analysis using repository data, consider whether you can apply the same processing pipeline to your own data.

Practical Workflow for Data Retrieval

A systematic workflow for retrieving genomic data from public repositories reduces errors and improves reproducibility. The steps below describe a general approach that applies across repositories.

Step 1: Define Your Data Requirements

Before searching any repository, write down the specific data characteristics you need. Include the organism, data type, sequencing platform, sample size, and any experimental conditions. For example, a study of endurance exercise training effects might require transcriptomic, proteomic, and metabolomic data from multiple tissues across multiple time points. The Molecular Transducers of Physical Activity Consortium generated a data compendium encompassing 9,466 assays across 19 tissues, 25 molecular platforms, and 4 training time points. Defining your requirements precisely helps you identify whether an existing dataset matches your needs.

Step 2: Search Across Relevant Repositories

Search multiple repositories instead of relying on a single source. The NCBI provides integrated search across its databases. The EBI offers complementary search tools and training resources. For cancer genomics, search the GDC directly. For proteomics, search iProX and other ProteomeXchange consortium members.

Use controlled vocabularies and ontology terms where available. The iProX repository deploys extensive controlled vocabularies and ontologies to annotate proteomics datasets. Searching with standardized terms improves recall and precision.

Step 3: Evaluate Access Requirements

Determine whether the datasets you need are open access or controlled access. Most non-human genomic data in public repositories are openly accessible. Human genomic data may require controlled access through data access committees or similar mechanisms. The NIH Genomic Data Sharing Policy describes the expectations for controlled access to certain human data.

Review the data use restrictions for each dataset. Some datasets may be limited to specific research areas, may prohibit commercial use, or may require acknowledgment of the data generators. The literature on public attitudes toward genomic data repositories indicates that clear governance and access policies are essential for maintaining trust. Respect the access conditions that apply to each dataset.

Step 4: Download and Verify Data Integrity

Download data using the repository's recommended transfer tools. Many repositories support high-speed transfer protocols such as Aspera. The iProX repository provides a fast Aspera-based transfer tool for proteomics data. Verify file integrity after download by checking checksums or other validation information provided by the repository.

Record the accession numbers, download dates, and file versions for every dataset you retrieve. This documentation supports reproducibility and allows you to cite the specific data versions used in your analysis.

Step 5: Document Your Retrieval Process

Maintain a retrieval log that records your search strategies, the repositories searched, the accession numbers retrieved, and any filtering decisions. This log serves as a record of your data provenance and supports the reproducibility of your analysis. When you publish results based on repository data, cite the specific accessions and versions you used.

Options and Tradeoffs Across Repositories

Different repositories offer different strengths and limitations. Understanding these tradeoffs helps you select the most appropriate repository for your specific needs.

NCBI Sequence Read Archive vs. European Nucleotide Archive

The NCBI Sequence Read Archive and the European Nucleotide Archive both store raw sequencing reads. The SRA is integrated with other NCBI resources, including GEO and the variation databases. The ENA is part of the EBI infrastructure and provides complementary services. For large-scale metagenomic analysis, the SRA holds over 600,000 publicly available metagenomes. The TOFU-MAaPO workflow can analyze metagenome files directly from the SRA using accession or study IDs, demonstrating the practical utility of SRA for large-scale projects.

Choose between SRA and ENA based on your familiarity with the respective interfaces, the geographic location of your computing resources, and the specific datasets you need. Both repositories accept data from researchers worldwide and make released data freely accessible.

NCBI Gene Expression Omnibus vs. Specialized Expression Databases

GEO serves as an international public repository for high-throughput microarray and next-generation sequence functional genomic data. It supports archiving of raw data, processed data, and metadata that are indexed, cross-linked, and searchable. GEO also provides web-based tools including GEO2R, an R-based web application that helps users analyze GEO data.

Specialized expression databases may offer more focused tools or curated datasets for specific organisms or conditions. However, GEO provides broad coverage and standardized data structures that facilitate cross-study comparisons. For most expression data retrieval needs, GEO is a reasonable first choice.

NCI Genomic Data Commons vs. General Repositories

The GDC provides cancer genomics data with harmonized processing across projects. This harmonization is a significant advantage for comparative analyses because it reduces technical variation between datasets. The GDC also provides cloud-enabled workflows that allow users to process their own data in the same manner as repository data.

General repositories such as SRA and GEO contain cancer genomics data but without the same level of harmonization. If your analysis requires comparing data across multiple cancer projects, the GDC offers substantial advantages. If you need raw data for custom processing, the SRA may be more appropriate.

Proteomics Repositories

The iProX repository serves as a public platform for collecting and sharing raw data, analysis results, and metadata from proteomics experiments. It is a full member of the ProteomeXchange consortium, and all released datasets are freely accessible to the public. iProX is based on a high availability architecture and has been deployed as part of the proteomics infrastructure of China.

For proteomics data, consider searching across all ProteomeXchange consortium members to maximize coverage. The consortium implements a mode of centralized metadata and distributed raw data management, which promotes effective data sharing.

Observations and Measurements in Repository Data

Genomic repositories contain data from diverse experimental designs and measurement platforms. Understanding the nature of these measurements is essential for appropriate data interpretation.

Sequencing Data Characteristics

Raw sequencing data vary in read length, depth, quality scores, and platform. For example, a transcriptomic study of apple peel tissue generated 28.5 million reads for control samples and 25.0 million reads for treated samples, corresponding to 2.9 Gb and 2.5 Gb of raw data. After quality filtering, the clean reads had GC content of approximately 47% and Q20 quality scores exceeding 99%. Approximately 95% of reads mapped to the reference genome.

When you retrieve sequencing data, examine the quality metrics reported by the data generators. Check read counts, quality score distributions, and mapping rates. These metrics affect the reliability of downstream analyses and the statistical power available for detecting biological differences.

Variant Data and Annotation

Variant repositories store information about genetic differences between individuals or populations. The IPD-IMGT/HLA Database contains over 25,000 allele sequences for 45 genes located within the major histocompatibility complex of the human genome. This region is the most polymorphic region of the human genome, with some genes having several thousand variants.

The database of genomic variants and related resources provide information about the clinical significance and population frequency of genetic variants. When you retrieve variant data, check the annotation standards used and the evidence supporting each variant classification.

Multi-Omic Data Integration

Some studies generate data across multiple molecular levels. The Molecular Transducers of Physical Activity Consortium profiled the temporal transcriptome, proteome, metabolome, lipidome, phosphoproteome, acetylproteome, ubiquitylproteome, epigenome, and immunome in whole blood, plasma, and 18 solid tissues. The resulting data compendium encompasses 9,466 assays across 19 tissues, 25 molecular platforms, and 4 training time points.

Multi-omic datasets enable integrated analyses that reveal relationships between molecular layers. However, they also present challenges for data retrieval and integration. When working with multi-omic data, verify that you retrieve all relevant data types and that sample identifiers are consistent across data layers.

Records and Documentation Practices

Maintaining accurate records of your repository use supports reproducibility and responsible data stewardship.

Accession Number Management

Record the accession numbers for all datasets you retrieve. Accession numbers are stable identifiers that allow others to locate the exact data you used. For sequencing data, record the study, sample, experiment, and run accessions as appropriate. For expression data, record the platform, sample, and series accessions.

Version Tracking

Repositories may update datasets over time. Record the version or release date of each dataset you retrieve. If a repository provides a data freeze or versioned release, use the version information in your records.

Data Use Agreements

For controlled access data, maintain copies of your data use agreements and any conditions attached to data access. These agreements may specify permitted uses, publication requirements, and restrictions on data sharing. Review these conditions before you use the data and before you share any derived results.

Quality Controls and Validation

Quality control is essential when working with repository data. The data you retrieve may contain errors, biases, or artifacts that affect your analysis.

Sequence Quality Assessment

Assess the quality of raw sequencing data before analysis. Check per-base quality scores, GC content, adapter contamination, and duplication rates. The apple transcriptomic dataset described above reported Q20 quality scores exceeding 99% for both control and treated samples, indicating high data quality. Lower quality data may require additional filtering or may be unsuitable for certain analyses.

Mapping and Alignment Validation

When you align sequencing reads to a reference genome, check mapping rates and coverage uniformity. The apple dataset reported approximately 95% of reads mapping to the reference genome. Low mapping rates may indicate contamination, reference mismatch, or data quality issues.

Cross-Platform Comparability

When comparing data from different platforms or laboratories, assess technical variation. The GDC harmonization approach addresses this challenge for cancer genomics data. For other data types, you may need to apply batch correction or other statistical methods to account for technical variation.

Common Failure Patterns in Repository Use

Researchers encounter recurring problems when using genomic data repositories. Recognizing these patterns helps you avoid them.

Incomplete Metadata

Datasets with incomplete metadata limit your ability to interpret results or reproduce analyses. Before downloading large files, verify that the metadata describes the experimental design, sample characteristics, and processing steps. If metadata are insufficient, consider whether the dataset is worth the storage and analysis effort.

Access Confusion

Researchers sometimes assume that all data in public repositories are openly accessible. Human genomic data may require controlled access under the NIH Genomic Data Sharing Policy or similar frameworks. Check access requirements before planning your analysis timeline.

Version Mismatch

Using different versions of reference genomes or annotation files across datasets can introduce errors. Verify that all data in your analysis use consistent reference versions. The GDC harmonization approach addresses this issue for cancer genomics data.

Overlooking Data Use Restrictions

Some datasets carry restrictions on use, including limitations on commercial applications or requirements for acknowledgment. Review data use conditions for each dataset and comply with all restrictions.

Download Interruptions

Large genomic datasets require substantial bandwidth and storage. Download interruptions can corrupt files or waste time. Use repository-recommended transfer tools and verify file integrity after download.

Limitations and Interpretation Boundaries

Genomic repository data have inherent limitations that affect interpretation.

Data Quality Heterogeneity

Repository data vary widely in quality. Some datasets were generated with rigorous quality controls, while others may contain errors or artifacts. The multi-batch reanalysis approach of jointly reevaluating gene and genome sequences from different works has gained particular relevance in the literature. However, this approach requires careful attention to data quality and processing consistency.

Population and Context Specificity

Data from one population or environment may not generalize to others. For example, a dataset of 352 nuclear genes from Rhododendron dauricum samples collected from seven distinct geographical populations in Northeast China provides high-resolution molecular markers for that species. These markers may not be transferable to other species or regions.

Technical Platform Effects

Different sequencing platforms and protocols introduce technical variation. The 16S rRNA gene sequencing workflow review highlights the impact of primer pair selection on diversity and taxonomic assignment outcomes in microbiome studies. When you compare data across studies, account for platform and protocol differences.

Annotation Currency

Genomic annotations change over time as new evidence accumulates. Variant classifications, gene models, and functional annotations may be updated. Check the annotation version used in your datasets and consider whether updated annotations would change your conclusions.

Safety and Regulatory Context

Genomic data use operates within a regulatory and ethical framework that researchers must respect.

Data Protection and Privacy

Human genomic data are sensitive personal information. The NIH Genomic Data Sharing Policy establishes expectations for protecting participant privacy while enabling data sharing for research. Controlled access mechanisms restrict data use to authorized researchers and specified purposes.

Commercial Use Considerations

Commercial interests in genomic data raise public concerns about consent, privacy, data security, trust, benefit sharing, and ownership. The literature on public attitudes toward biobank and genomic data repositories indicates that carefully considered regulatory and data governance and access policies are required to maintain public trust. When you use genomic data, consider whether your intended use aligns with the expectations of data contributors and the public.

International Data Transfer

Genomic data may be subject to international data transfer regulations. Researchers working across national boundaries should verify that their data use complies with applicable laws and policies. The iProX repository, for example, is deployed as part of the proteomics infrastructure of China and follows the ProteomeXchange consortium standards.

Professional Escalation Criteria

Some situations require consultation with experts or escalation to repository administrators.

When to Seek Expert Assistance

Consult a bioinformatics specialist or repository help desk when you encounter any of the following situations:

  • You cannot locate data that should exist based on a publication or prior knowledge
  • You encounter access errors or permission issues that you cannot resolve
  • You need to retrieve data from a repository with unfamiliar data structures
  • You plan a large-scale retrieval that requires substantial computing resources
  • You need to harmonize data across repositories with different processing pipelines

When to Contact Repository Administrators

Contact repository administrators when you identify data errors, metadata problems, or access issues. Repositories rely on user feedback to improve data quality. Report any discrepancies between publications and repository records.

When to Consult Ethics or Compliance Officers

Consult ethics or compliance officers when your intended data use raises questions about consent, privacy, or data use restrictions. This consultation is particularly important for human genomic data and for data with commercial implications.

Frequently Asked Questions

What is the difference between the NCBI Sequence Read Archive and the Gene Expression Omnibus?

The Sequence Read Archive stores raw sequencing reads from DNA and RNA sequencing experiments. The Gene Expression Omnibus stores functional genomics data including microarray and high-throughput sequencing expression data, along with processed data and metadata. If you need raw sequence reads for variant calling or assembly, use SRA. If you need expression measurements for differential expression analysis, use GEO.

How do I access controlled access data in the NCI Genomic Data Commons?

Controlled access data in the GDC require authorization through the appropriate data access mechanism. Review the access requirements for each dataset and apply through the designated process. The NIH Genomic Data Sharing Policy describes the expectations for controlled access to certain human genomic data.

What is the ProteomeXchange consortium and how does it relate to iProX?

The ProteomeXchange consortium implements a mode of centralized metadata and distributed raw data management to promote effective data sharing for proteomics. iProX is a full member of the consortium and serves as a public platform for collecting and sharing raw data, analysis results, and metadata from proteomics experiments. All released datasets in iProX are freely accessible to the public.

How can I ensure that my analysis of repository data is reproducible?

Document your retrieval process, record accession numbers and versions, and maintain a retrieval log. Apply consistent processing pipelines to all datasets in your analysis. The GDC provides executable workflows that can be applied to reproducibly process user data and harmonize datasets across repositories.

What should I do if a dataset lacks sufficient metadata for my analysis?

Contact the data submitter if contact information is available. Consider whether the dataset is suitable for your purpose without the missing metadata. The Genomic Observatories Metadatabase was developed to address the need for field and sampling event metadata associated with genetic samples, indicating the importance of metadata for data reuse.

Are there restrictions on commercial use of genomic data from public repositories?

Some datasets carry restrictions on commercial use. Review the data use conditions for each dataset before use. The literature on public attitudes toward genomic data repositories indicates that commercial involvement raises concerns about consent, privacy, benefit sharing, and ownership. Respect all data use restrictions.

How do I choose between the NCBI and EBI repositories for my data retrieval?

Consider your familiarity with the respective interfaces, the geographic location of your computing resources, and the specific datasets you need. Both NCBI and EBI provide comprehensive genomic data resources. The EBI provides training materials that can help you learn their systems.

What is the FAIR Guiding Principles framework and why does it matter for repository use?

The FAIR Guiding Principles state that data should be Findable, Accessible, Interoperable, and Reusable. These principles provide a framework for evaluating repository quality and data management practices. When you select a repository for data retrieval or deposition, assess whether the repository supports persistent identifiers, clear access conditions, standardized metadata, and explicit reuse licenses.

Related Bioinformatics Guides

References and Further Reading

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.