NCBI GEO vs. ArrayExpress vs. ENA: Which RNA-seq Repository Should You Use?
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- For depositing raw sequencing reads, the European Nucleotide Archive (ENA) is the designated repository for European-funded projects and accepts data globally, while the NCBI Sequence Read Archive (SRA) serves a similar role for US-funded projects.
- For processed expression matrices and rich sample annotations, NCBI Gene Expression Omnibus (GEO) is the most widely adopted repository in published literature, particularly for US-based research and international meta-analyses.
- ArrayExpress, now integrated with the BioStudies database for new submissions, historically served as the European counterpart to GEO for processed functional genomics data, and remains accessible for archived experiments.
- A common and recommended practice for RNA-seq projects is dual deposition: raw reads in ENA/SRA and processed count matrices with metadata in GEO/ArrayExpress to satisfy journal and funding mandates.
- Data discoverability and reanalysis are significantly impacted by metadata quality; ENA uses the INSDC standard, GEO employs MINiML/SOFT formats, and ArrayExpress/BioStudies utilize MAGE-TAB and the BioStudies annotation model, respectively.
- Researchers should consult funding agreements and target journals to determine specific repository mandates, as European Commission projects often require ENA/ArrayExpress, while NIH projects typically mandate SRA/GEO.
Researchers generating RNA-seq data face a practical decision early in their project: where to deposit raw sequencing reads, processed count matrices, and associated metadata. The choice affects how easily collaborators find your data, how smoothly journals accept your submission, and how efficiently you can retrieve comparable public datasets for reanalysis. This article compares the three major public repositories that host RNA-seq data, the National Center for Biotechnology Information Gene Expression Omnibus (NCBI GEO), the European Bioinformatics Institute ArrayExpress, and the European Nucleotide Archive (ENA), with attention to data coverage, submission requirements, search capabilities, and data formats.
The direct answer to which repository you should use depends on your primary need. For depositing raw sequencing reads, ENA is the designated archive for European-funded projects and accepts data from anywhere. For depositing processed expression matrices with rich sample annotations, GEO is the most widely used choice in the published literature. For retrieving curated functional genomics datasets, GEO and ArrayExpress both offer structured query interfaces, but GEO currently hosts a substantially larger collection of RNA-seq series. Many projects deposit raw reads in ENA and processed counts in GEO simultaneously, and this dual-deposit pattern is common in the literature.
Scope and Reader Context
This comparison serves biology students, researchers, laboratory professionals, and life-science practitioners who need to decide where to submit new RNA-seq data or where to search for existing datasets. The practical outcome is a decision framework based on your funding source, target journal, data type, and analysis workflow. The article covers the submission process for each repository, the search interfaces available for data retrieval, the file formats each repository accepts and distributes, and the metadata standards that govern discoverability.
The comparison does not cover every bioinformatics resource hosted by NCBI or EMBL-EBI. It focuses on the three repositories most relevant to RNA-seq workflows. Related resources such as the NCBI Sequence Read Archive (SRA) and the EMBL-EBI BioStudies database are discussed where they interact with the primary repositories.
At a Glance
| Feature | NCBI GEO | ArrayExpress | ENA |
|---|---|---|---|
| Primary data type | Processed expression matrices, microarray and RNA-seq series | Processed functional genomics data, microarray and RNA-seq series | Raw sequencing reads, assembled sequences, and associated metadata |
| Submission route | GEO submission portal with SOFT, MINiML, or spreadsheet formats | BioStudies submission pipeline with ArrayExpress annotation spreadsheets | Webin submission system with ERA, SRA, or ENA file formats |
| Search interface | GEO DataSets and GEO Profiles browsers with Entrez query system | Bioconductor ArrayExpress package and EBI search portal | ENA browser with text search and programmatic access via ENA API |
| File formats accepted | Processed data tables, raw data optional for microarray, raw sequence files routed to SRA | Processed data files, raw data routed to ENA | FASTQ, BAM, CRAM, and other sequence formats |
| Metadata standard | MINiML and SOFT formats with GEO sample and series attributes | MAGE-TAB and BioStudies annotation model | INSDC standard with sample, experiment, run, and analysis objects |
| Typical use in RNA-seq | Differential expression analysis, meta-analysis, dataset discovery | Functional genomics reanalysis, cross-species comparisons | Read alignment, variant calling, transcript assembly, data archiving |
| Accession prefix | GSE for series, GSM for samples, GDS for datasets | E-MTAB for ArrayExpress experiments | ERP for projects, ERX for experiments, ERR for runs |
| Funding mandate alignment | NIH and US-based grant requirements | European Commission and UKRI requirements | European Commission and global INSDC membership |
Repository Roles in the RNA-seq Data Lifecycle
RNA-seq data exists in multiple forms throughout a project. Raw sequencing reads are produced by the instrument as FASTQ files. These reads are aligned to a reference genome or transcriptome, producing BAM files. Quantification tools generate count matrices that summarize expression levels for each gene or transcript. Downstream analysis produces differential expression tables, pathway enrichments, and visualizations.
The three repositories occupy different positions in this lifecycle. ENA archives raw sequencing reads and serves as the primary repository for sequence data in Europe. GEO archives processed expression data, including count matrices and clinical or experimental annotations, and is the dominant repository for functional genomics data in the United States. ArrayExpress historically served a similar role to GEO in Europe but has transitioned its submission pipeline to the BioStudies database while maintaining access to archived functional genomics datasets.
A researcher depositing RNA-seq data must decide whether to submit raw reads, processed counts, or both. Journals increasingly require raw sequencing data deposition as a condition of publication. Funding agencies often mandate deposition in specific repositories. The practical answer for most projects is to deposit raw reads in ENA or SRA and processed expression matrices in GEO or ArrayExpress.
NCBI GEO: The Dominant Repository for Processed Expression Data
The National Center for Biotechnology Information hosts the Gene Expression Omnibus as a public functional genomics data repository. GEO accepts array-based and high-throughput sequencing data related to gene expression, and it supports tools for querying and downloading experiments and curated gene expression profiles. The NCBI platform provides integrated search across GEO DataSets, GEO Profiles, and related sequence resources, which makes it a practical first stop for researchers seeking public RNA-seq data.
Data Coverage and Growth
GEO contains a large and growing collection of RNA-seq series. The repository has operated since the early 2000s and has accumulated submissions from diverse organisms, experimental designs, and disease contexts. Published studies frequently cite GEO accession numbers, and the repository serves as the default destination for processed expression data in the United States and internationally.
The volume of RNA-seq data in GEO reflects the broader adoption of the technology. Many studies deposit both raw sequence data in SRA and processed count matrices in GEO. This dual-deposit pattern means that a researcher searching GEO for a disease of interest will find series with sample annotations, experimental design details, and processed data tables that can be downloaded directly for reanalysis.
Submission Requirements and Process
GEO submission requires a structured set of metadata for each sample and series. The submission portal accepts data in several formats, including SOFT files, MINiML files, and Excel spreadsheets. For RNA-seq data, the processed count matrix is the primary data file, while raw sequence reads are routed to SRA as part of the submission process.
The GEO submission process requires the following components:
- A series title and summary describing the experiment
- Overall design information explaining the experimental groups and comparisons
- Sample metadata including organism, tissue, cell type, treatment, and any relevant clinical variables
- Processed data files containing expression measurements for each sample
- Platform information describing the sequencing instrument and alignment approach
The submission system validates metadata fields and flags missing or inconsistent information. Submitters can update their records after deposition, which is useful when additional analyses are completed or when reviewers request additional metadata.
Search Capabilities and Data Retrieval
GEO offers multiple search interfaces. The GEO DataSets browser allows users to query by organism, platform, study type, and free-text terms. The GEO Profiles browser provides gene-level expression views across multiple datasets. The Entrez system integrates GEO with other NCBI databases, allowing users to move from a gene record to expression data for that gene across all GEO series.
For programmatic access, the NCBI provides the GEOquery package through Bioconductor. This package allows researchers to download GEO series directly into R for analysis. The Bioconductor project maintains documentation and workflows for using GEOquery with RNA-seq data, including steps for reading count matrices, accessing sample metadata, and preparing data for differential expression analysis.
Practical Use Cases for GEO
Researchers commonly use GEO for three purposes. First, they search for public datasets relevant to their biological question, download the processed count matrices, and perform their own differential expression analysis. Second, they use GEO series as validation cohorts to confirm findings from their own experiments. Third, they aggregate multiple GEO datasets for meta-analysis, combining expression data across studies to increase statistical power.
The published literature demonstrates these use patterns. Studies of bladder cancer have used gene and miRNA expression data from NCBI GEO to identify differentially expressed genes and build survival models. Research on Kawasaki disease has retrieved mRNA expression profiles from GEO to investigate the role of specific gene families. Studies of gastric mucosa after bariatric surgery have downloaded GEO series to identify differentially expressed genes and pathway enrichments.
ArrayExpress: The European Functional Genomics Archive
ArrayExpress is the functional genomics data repository at the European Bioinformatics Institute. It accepts submissions of microarray and high-throughput sequencing data and provides search and download services for processed expression data. ArrayExpress has historically served as the European counterpart to GEO, and many European journals and funding agencies have required or recommended deposition in this repository.
Current Status and Transition to BioStudies
ArrayExpress has undergone a significant transition in recent years. The submission pipeline for new functional genomics datasets has moved to the BioStudies database, which provides a broader framework for linking data files, protocols, and associated publications. Existing ArrayExpress datasets remain accessible, and the search interface continues to provide access to archived experiments.
This transition affects researchers planning new submissions. instead of submitting directly to ArrayExpress, new functional genomics datasets are submitted through BioStudies, with processed data files and metadata organized according to the BioStudies annotation model. The ArrayExpress accession format, such as E-MTAB, continues to be used for experiments archived in the system.
Data Coverage and Relationship to ENA
ArrayExpress contains a substantial collection of RNA-seq experiments, particularly from European research groups. The repository includes both microarray and high-throughput sequencing data, with processed expression matrices and sample annotations. For RNA-seq submissions, raw sequence reads are deposited in ENA, while processed data and experimental metadata are maintained in ArrayExpress or BioStudies.
The relationship between ArrayExpress and ENA mirrors the relationship between GEO and SRA. The processed data repository handles expression matrices and annotations, while the sequence archive handles raw reads. Researchers retrieving data from ArrayExpress receive processed count matrices, while researchers needing raw reads must access ENA.
Submission Requirements and Process
New submissions to ArrayExpress follow the BioStudies submission workflow. The process requires:
- A study description with title, abstract, and experimental design
- Sample annotations using controlled vocabulary where available
- Processed data files in tabular or matrix format
- Protocol descriptions for library preparation, sequencing, and data processing
- Links to raw sequence data deposited in ENA
The submission system uses annotation spreadsheets to collect sample metadata. These spreadsheets follow a structured format with columns for sample characteristics, experimental factors, and technical variables. The system validates the spreadsheets and provides feedback on missing or malformed entries.
Search Capabilities and Data Retrieval
ArrayExpress provides a web search interface that supports queries by organism, experiment type, and free-text terms. The EBI search portal integrates ArrayExpress with other EMBL-EBI resources, allowing users to search across multiple databases from a single query.
For programmatic access, the Bioconductor ArrayExpress package provides tools for downloading experiments directly into R. This package supports searching for experiments by accession or keywords, downloading processed data files, and reading sample metadata into R data structures. The package documentation includes examples of retrieving RNA-seq experiments and preparing data for analysis.
Practical Use Cases for ArrayExpress
Researchers use ArrayExpress for similar purposes as GEO, including dataset discovery, validation, and meta-analysis. The repository is particularly relevant for researchers working with European collaborators or submitting to European journals. The EMBL-EBI training program offers courses and materials on using ArrayExpress and related resources for functional genomics analysis.
The published literature includes examples of ArrayExpress data use. Systems biology studies of cystic fibrosis have built gene regulatory networks from publicly available transcriptomic datasets, including those archived in European repositories. Studies of alternative splicing have used RNA-seq data from public repositories to characterize isoform diversity across organisms.
ENA: The European Nucleotide Archive for Raw Sequencing Reads
The European Nucleotide Archive is the European component of the International Nucleotide Sequence Database Collaboration, which also includes the NCBI Sequence Read Archive and the DNA Data Bank of Japan. ENA accepts raw sequencing reads, assembled sequences, and associated metadata, and it provides search and download services for sequence data.
Data Coverage and Role in the INSDC
ENA is the designated repository for raw sequencing data from European-funded projects. The archive accepts submissions from researchers worldwide and provides permanent accession numbers for sequence data. ENA is part of the INSDC, which means that data submitted to ENA is synchronized with SRA and DDBJ, allowing researchers to access the same data through any of the three archives.
For RNA-seq projects, ENA stores FASTQ files containing raw sequencing reads, along with metadata describing the sequencing instrument, library preparation, and sample characteristics. The archive also stores BAM and CRAM files for aligned reads, and it provides access to assembled transcript sequences when relevant.
Submission Requirements and Process
ENA submissions use the Webin submission system. The process requires:
- A study or project registration describing the overall research project
- Sample registrations with organism and sample attributes
- Experiment registrations describing the sequencing library and instrument
- Run registrations linking sequence files to experiments
- Sequence file uploads in FASTQ, BAM, CRAM, or other accepted formats
The Webin system validates file formats and metadata fields, and it provides accession numbers at each stage of the submission. The system supports both interactive submissions through the web interface and programmatic submissions through the Webin REST API.
Search Capabilities and Data Retrieval
ENA provides a browser interface that supports text searches and filtering by organism, study type, and other attributes. The ENA API allows programmatic access to sequence data and metadata, supporting automated retrieval of large datasets.
For RNA-seq data, the ENA browser provides access to run files containing raw reads. Researchers can download FASTQ files for alignment and quantification, or they can access BAM files for aligned reads. The ENA API supports downloading files in bulk, which is useful for large-scale reanalysis projects.
Practical Use Cases for ENA
Researchers use ENA for three primary purposes. First, they deposit raw sequencing reads to satisfy journal and funding agency requirements. Second, they retrieve raw reads from public projects to perform their own alignment and quantification. Third, they access assembled sequences for transcript characterization.
The published literature demonstrates the importance of raw sequence data access. Studies of non-triplet alternative splicing have used RNA-seq data to characterize isoform diversity, requiring access to raw reads for alignment and quantification. Studies of cardiac fibrosis have conducted meta-analyses of public single-cell sequencing datasets, requiring retrieval of raw or processed data from multiple repositories.
Comparing Submission Workflows for RNA-seq Projects
The submission workflow for an RNA-seq project depends on which repository you choose and what data types you need to deposit. The following comparison covers the practical steps for each repository.
GEO Submission Workflow
Submitting processed RNA-seq data to GEO involves the following steps:
- Prepare processed count matrices for each sample, with genes as rows and samples as columns
- Prepare sample metadata, including organism, tissue, treatment, and experimental group
- Create a GEO submission using the web portal or spreadsheet templates
- Upload processed data files and metadata
- Submit raw sequence reads to SRA if required by the journal or funding agency
- Obtain GEO accession numbers for the series and samples
- Cite the GEO accession in your manuscript
The GEO submission system provides validation feedback during the submission process. Common issues include missing metadata fields, inconsistent sample names, and malformed data files. The system allows submitters to revise records after deposition.
ArrayExpress and BioStudies Submission Workflow
Submitting processed RNA-seq data through the BioStudies pipeline involves the following steps:
- Prepare processed count matrices and sample annotation spreadsheets
- Register the study in BioStudies and obtain a study accession
- Upload processed data files and annotation spreadsheets
- Submit raw sequence reads to ENA using the Webin system
- Link the ENA study accession to the BioStudies record
- Obtain the ArrayExpress or BioStudies accession for citation
The BioStudies submission system uses annotation spreadsheets that follow a structured format. The system validates the spreadsheets and provides feedback on missing or inconsistent entries. The submission process requires careful attention to controlled vocabulary terms for sample characteristics.
ENA Submission Workflow
Submitting raw sequencing reads to ENA involves the following steps:
- Register a study in the Webin system and obtain a project accession
- Register samples with organism and sample attributes
- Register experiments describing the sequencing library and instrument
- Upload sequence files in FASTQ, BAM, or CRAM format
- Obtain run accessions for each sequence file
- Link the ENA study accession to any processed data deposited in GEO or ArrayExpress
The Webin system provides validation of file formats and metadata. The system supports both interactive and programmatic submissions, and it provides accession numbers at each stage.
Search and Retrieval Comparison for RNA-seq Data
Finding existing RNA-seq data requires different search strategies for each repository. The following comparison covers the search interfaces and retrieval methods available.
Searching GEO for RNA-seq Data
The GEO DataSets browser supports queries by organism, study type, platform, and free-text terms. A typical search for RNA-seq data involves the following steps:
- Navigate to the GEO DataSets browser
- Enter search terms for the biological condition of interest
- Filter by organism and study type
- Review the list of matching series
- Open series records to review sample metadata and experimental design
- Download processed data files or use GEOquery to load data into R
The GEO Profiles browser provides gene-level views across datasets. This interface is useful for checking expression patterns of specific genes across multiple studies.
Searching ArrayExpress for RNA-seq Data
The ArrayExpress search interface supports queries by organism, experiment type, and free-text terms. The EBI search portal integrates ArrayExpress with other EMBL-EBI resources. A typical search involves the following steps:
- Navigate to the ArrayExpress search interface or the EBI search portal
- Enter search terms for the biological condition of interest
- Filter by organism and experiment type
- Review the list of matching experiments
- Open experiment records to review sample metadata and protocols
- Download processed data files or use the ArrayExpress Bioconductor package
Searching ENA for RNA-seq Data
The ENA browser supports text searches and filtering by organism, study type, and other attributes. The ENA API provides programmatic access. A typical search involves the following steps:
- Navigate to the ENA browser
- Enter search terms for the biological condition of interest
- Filter by organism and data type
- Review the list of matching studies and runs
- Open run records to review sequencing metadata
- Download FASTQ files or use the ENA API for bulk retrieval
Data Formats and Compatibility with Analysis Workflows
The format of data retrieved from each repository affects how easily it can be used in downstream analysis workflows. The following comparison covers the formats available from each repository.
Processed Data Formats
GEO provides processed data in SOFT and MINiML formats, along with supplementary files that often contain count matrices in tabular format. The GEOquery Bioconductor package reads these formats and converts them into R data structures suitable for analysis.
ArrayExpress provides processed data in MAGE-TAB format, along with supplementary files containing expression matrices. The ArrayExpress Bioconductor package reads these formats and loads data into R.
ENA provides raw sequence data in FASTQ, BAM, and CRAM formats. These formats require alignment and quantification before expression analysis can proceed.
Raw Sequence Data Formats
ENA and SRA store raw sequencing reads in FASTQ format, along with aligned reads in BAM or CRAM format. These files are typically large, ranging from hundreds of megabytes to hundreds of gigabytes per sample.
The choice between processed and raw data affects the analysis workflow. Processed count matrices can be used directly for differential expression analysis, while raw reads require alignment and quantification steps. The Galaxy Training Network provides tutorials for RNA-seq analysis workflows that start from raw reads and produce count matrices. The nf-core project provides community-developed pipelines for RNA-seq analysis that accept raw reads as input and produce count matrices and quality reports.
Metadata Formats and Standards
The metadata standards used by each repository affect data discoverability and interoperability. GEO uses the MINiML and SOFT formats, which encode sample attributes, experimental design, and data processing information. ArrayExpress uses the MAGE-TAB format, which provides a structured representation of experimental metadata. ENA uses the INSDC standard, which defines sample, experiment, run, and analysis objects.
The FAIR guiding principles for data management and reuse emphasize that data should be findable, accessible, interoperable, and reusable. The metadata standards used by these repositories support these principles by providing structured descriptions of experimental context. Studies that reuse public omics data have demonstrated the value of well-annotated repositories for enabling new research findings.
Practical Implementation Steps for Repository Selection
Selecting the appropriate repository for your RNA-seq project requires a systematic assessment of your requirements. The following steps provide a practical framework for this decision.
Step 1: Review Funding and Journal Requirements
Check your funding agreement and target journal for data deposition requirements. European Commission funded projects typically require deposition in ENA for raw reads and ArrayExpress or BioStudies for processed data. NIH funded projects typically require deposition in SRA for raw reads and GEO for processed data. Many journals require raw data deposition as a condition of publication.
Step 2: Determine Your Data Types
Identify whether your project generates raw sequencing reads, processed count matrices, or both. Raw reads require deposition in ENA or SRA. Processed count matrices require deposition in GEO or ArrayExpress. Most RNA-seq projects generate both data types and require dual deposition.
Step 3: Assess Metadata Requirements
Review the metadata requirements for each repository. GEO requires detailed sample annotations and experimental design information. ArrayExpress and BioStudies require annotation spreadsheets with controlled vocabulary terms. ENA requires sample attributes and sequencing metadata. Choose the repository where you can provide the required metadata without excessive burden.
Step 4: Evaluate Search and Retrieval Needs
Consider whether you need to retrieve data from the repository for your own analysis. GEO provides the most mature search interface for processed expression data. ENA provides the most comprehensive access to raw sequence data. ArrayExpress provides access to European functional genomics datasets.
Step 5: Plan for Dual Deposition
For most RNA-seq projects, plan to deposit raw reads in ENA or SRA and processed data in GEO or ArrayExpress. This dual-deposit pattern satisfies journal and funding requirements while maximizing data discoverability. Link the raw and processed data records through accession numbers.
Records and Measurements for Repository Selection
Maintaining records of your repository selection and submission process supports reproducibility and compliance. The following records are recommended for RNA-seq projects.
Submission Records
Document the following information for each repository submission:
- Repository name and accession numbers
- Submission date and submission system used
- Data files uploaded and their formats
- Metadata fields provided and any controlled vocabulary terms used
- Validation feedback received and corrections made
- Links between raw and processed data records
Data Retrieval Records
Document the following information for each data retrieval:
- Repository name and accession numbers of retrieved datasets
- Search terms and filters used
- Download date and file formats retrieved
- Software versions used for data processing
- Any data quality issues observed
Analysis Records
Document the following information for each analysis using public data:
- Source repository and accession numbers
- Data processing steps and software versions
- Quality control metrics and filtering decisions
- Analysis parameters and statistical methods
- Output files and their locations
Common Failure Patterns in Repository Selection and Use
Researchers encounter several common problems when selecting and using RNA-seq repositories. Recognizing these patterns helps avoid costly delays and data loss.
Failure Pattern 1: Submitting Only Processed Data
Some researchers deposit only processed count matrices in GEO or ArrayExpress and skip raw read deposition. This pattern fails to satisfy journal and funding requirements that mandate raw data availability. Reviewers may request raw reads for independent verification of results. The solution is to deposit raw reads in ENA or SRA alongside processed data.
Failure Pattern 2: Submitting Only Raw Reads
The opposite pattern involves depositing only raw reads in ENA or SRA without processed count matrices. This pattern makes it difficult for other researchers to reuse the data, because they must perform their own alignment and quantification before they can interpret the results. The solution is to deposit processed count matrices in GEO or ArrayExpress with appropriate sample annotations.
Failure Pattern 3: Incomplete Metadata
Submissions with incomplete or inconsistent metadata are difficult to discover and reuse. Missing sample attributes, unclear experimental design descriptions, and inconsistent sample names all reduce data utility. The solution is to follow the metadata standards for each repository and to provide comprehensive annotations for every sample.
Failure Pattern 4: Ignoring Repository-Specific Requirements
Each repository has specific submission requirements and file format expectations. Submissions that ignore these requirements are rejected or delayed. The solution is to review the submission documentation for each repository before starting the submission process.
Failure Pattern 5: Failing to Link Raw and Processed Data
Dual depositions that do not link raw and processed data records create confusion for data reusers. Researchers may find the processed data but not know where to access the raw reads, or they may find the raw reads but not know where to access the processed counts. The solution is to include cross-references between repository records.
Quality Controls for Repository Data
Quality control is essential when using public RNA-seq data from any repository. The following controls help ensure that retrieved data is suitable for analysis.
Metadata Quality Assessment
Review the metadata for each dataset before downloading data. Check that sample annotations are complete and consistent with the experimental design. Look for missing values, inconsistent terminology, and unclear descriptions. Datasets with poor metadata may not be suitable for your analysis.
Data Quality Assessment
Assess the quality of processed count matrices before analysis. Check for missing values, zero counts, and unusual distributions. Compare sample-level metrics such as total read counts and gene detection rates. The SkewC tool provides a method for identifying cells with skewed gene body coverage in single-cell RNA-seq data, which is relevant for quality assessment in scRNA-seq workflows.
Reproducibility Assessment
Check whether the dataset includes sufficient information to reproduce the analysis. Look for descriptions of alignment methods, quantification tools, and filtering criteria. Datasets that lack processing details may not be suitable for meta-analysis or validation studies.
Limitations of Repository Data
Public RNA-seq data from any repository has limitations that affect its utility for reanalysis. Understanding these limitations helps researchers interpret results appropriately.
Batch Effects and Technical Variation
RNA-seq data generated in different laboratories, at different times, and with different protocols exhibit technical variation that can confound biological comparisons. Meta-analyses that combine data from multiple studies must account for batch effects through statistical methods or study design.
Incomplete Metadata
Many public datasets lack complete metadata, making it difficult to interpret experimental conditions or to identify appropriate samples for comparison. Researchers must carefully review available metadata and document any assumptions made during data selection.
Processing Pipeline Differences
Different studies use different alignment and quantification pipelines, producing count matrices that are not directly comparable. Researchers must decide whether to use processed data as provided or to reprocess raw reads with a consistent pipeline.
Sample Size Limitations
Public datasets often have small sample sizes, particularly for rare diseases or specialized experimental conditions. Small sample sizes limit statistical power and increase the risk of false discoveries.
Safety and Regulatory Context
Data deposition in public repositories has regulatory and ethical implications that researchers must consider.
Data Privacy and Consent
RNA-seq data derived from human samples must comply with privacy regulations and consent requirements. Researchers must ensure that deposited data does not contain identifiable information and that the original consent covers public data sharing. Some repositories provide controlled access options for sensitive data.
Data Use Agreements
Some datasets are subject to data use agreements that restrict how the data can be used. Researchers retrieving data from public repositories must review and comply with any applicable data use agreements.
Export Controls
Raw sequencing data may be subject to export controls in some jurisdictions. Researchers depositing or retrieving data across international borders must comply with applicable regulations.
Professional Escalation Criteria
Researchers should seek professional guidance when encountering specific situations during repository selection or data use.
When to Consult a Bioinformatics Specialist
Consult a bioinformatics specialist when you need to:
- Choose between processed and raw data for your analysis
- Design a reprocessing pipeline for raw reads from public datasets
- Integrate data from multiple repositories with different formats
- Address batch effects in meta-analyses of public data
When to Consult a Data Manager or Librarian
Consult a data manager or librarian when you need to:
- Understand funding agency data deposition requirements
- Prepare metadata for repository submission
- Navigate controlled access data request processes
- Comply with data privacy regulations
When to Consult an Ethics or Regulatory Advisor
Consult an ethics or regulatory advisor when you need to:
- Determine whether your data deposition plan complies with consent requirements
- Assess whether your data contains identifiable information
- Navigate international data transfer regulations
- Address data use agreement restrictions
Frequently Asked Questions
What is the main difference between GEO and ENA for RNA-seq data?
GEO stores processed expression data, including count matrices and sample annotations, while ENA stores raw sequencing reads in FASTQ, BAM, or CRAM format. For a complete RNA-seq submission, you typically deposit raw reads in ENA and processed counts in GEO. The two repositories serve complementary roles in the data lifecycle.
Can I submit the same RNA-seq dataset to both GEO and ArrayExpress?
You can deposit processed data in both GEO and ArrayExpress, but this is rarely necessary. Most projects deposit processed data in one repository and raw reads in the corresponding sequence archive. The choice between GEO and ArrayExpress typically depends on your funding source and target journal.
Do I need to deposit raw sequencing reads if I already deposited processed counts in GEO?
Most journals and funding agencies require raw sequencing read deposition in addition to processed counts. Raw reads allow other researchers to independently verify your analysis and to apply alternative processing methods. Check your target journal and funding agreement for specific requirements.
How do I search for RNA-seq datasets in GEO?
Use the GEO DataSets browser to search by organism, study type, and free-text terms. Filter results by platform and study design. Open series records to review sample metadata and download processed data files. The GEOquery Bioconductor package provides programmatic access for downloading data into R.
What file formats does ENA accept for RNA-seq data?
ENA accepts FASTQ files for raw sequencing reads, BAM and CRAM files for aligned reads, and FASTA files for assembled sequences. The Webin submission system validates file formats and provides accession numbers for each submission stage.
How long does it take to submit RNA-seq data to these repositories?
Submission time varies depending on data size and metadata complexity. Processed data submissions to GEO or ArrayExpress can be completed in hours if metadata is prepared in advance. Raw read submissions to ENA can take longer due to file upload times for large datasets.
Can I access data from these repositories programmatically?
Yes, all three repositories support programmatic access. GEO provides the GEOquery Bioconductor package and the Entrez API. ArrayExpress provides the ArrayExpress Bioconductor package and the EBI search API. ENA provides the ENA API for bulk data retrieval.
What should I do if I find errors in a public dataset?
Contact the data submitter through the repository contact information to report errors. Many submitters appreciate feedback and may update their records. If the errors affect your analysis, document the issues and consider excluding the dataset from your analysis.
Related Bioinformatics Guides
- RNA-Seq Databases: Accessing and Using Public RNA-Seq Data
- RNA-Seq Data Analysis in Galaxy: A User-Friendly Platform
- RNA-Seq Data Analysis Workflow: From Raw Reads to Insights
- Genomic Data Repositories: Navigating Public Databases for Research
- RNA-Seq vs qPCR: Validation and Comparison
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
- NCBI Data Resources. National Center for Biotechnology Information.
- EMBL-EBI Training. European Bioinformatics Institute.
- Bioconductor. Bioconductor Project.
- Galaxy Training Network. Galaxy Project.
- nf-core Documentation. nf-core.
- The Carpentries Lessons. The Carpentries.
- Identification of biomedical entities from multiple repositories using a specialized metadata schema and search-augmented large language models.. 2026.
- Exploiting open source omics data to advance pancreas research.. 2024.
- Pervasive non-triplet alternative splicing drives functional isoform diversity.. 2026.
- A personalized medicine approach identifies enasidenib as an efficient treatment for IDH2 mutant chondrosarcoma.. 2024.
- From CFTR to a CF signalling network: a systems biology approach to study Cystic Fibrosis.. 2024.
- Selective inhibition of stromal mechanosensing suppresses cardiac fibrosis.. 2025.
- SkewC: Identifying cells with skewed gene body coverage in single-cell RNA sequencing data.. 2022.
- Age-Stratified Glycolysis and Bile Acid Metabolism Drive Survival in Acute Liver Failure: Insights from NCBI Pathways. Research Review, 2025.
- An unfolded protein response (UPR)-signature regulated by the NFKB-miR-29b/c axis fosters tumor aggressiveness and poor survival in bladder cancer. Frontiers in Molecular Biosciences, 2025.
- Hub genes and key pathways of Graves’ disease: bioinformatics analysis and validation. HORMONES, 2025.
- Algorithm-based assessment of T-cell dysfunction and exclusion to forecast ICB sensitivity in pediatric brain ependymoma. Journal of Neuro-Oncology, 2025.
- Gastric mucosal differentially expressed genes after bariatric surgery: Effects on sterol-related pathways.. Journal of Steroid Biochemistry and Molecular Biology, 2025.
- Integrative Machine Learning Framework for Effector Gene Prediction in Magnaporthe Oryzae. Educational Sciences International Conference, 2026.
- Argonaute2 and Argonaute4 Involved in the Pathogenesis of Kawasaki Disease via mRNA Expression Profiles. Children, 2025.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.