NCBI GEO vs. ArrayExpress vs. ENA: Which RNA-seq Repository Should You Use?

By Dr. Zubair Khalid, DVM, MS, PhD ·

NCBI GEO vs. ArrayExpress vs. ENA: Which RNA-seq Repository Should You Use?

Key Takeaways

  • For depositing raw sequencing reads, the European Nucleotide Archive (ENA) is the designated repository for European-funded projects and accepts data globally, while the NCBI Sequence Read Archive (SRA) serves a similar role for US-funded projects.
  • For processed expression matrices and rich sample annotations, NCBI Gene Expression Omnibus (GEO) is the most widely adopted repository in published literature, particularly for US-based research and international meta-analyses.
  • ArrayExpress, now integrated with the BioStudies database for new submissions, historically served as the European counterpart to GEO for processed functional genomics data, and remains accessible for archived experiments.
  • A common and recommended practice for RNA-seq projects is dual deposition: raw reads in ENA/SRA and processed count matrices with metadata in GEO/ArrayExpress to satisfy journal and funding mandates.
  • Data discoverability and reanalysis are significantly impacted by metadata quality; ENA uses the INSDC standard, GEO employs MINiML/SOFT formats, and ArrayExpress/BioStudies utilize MAGE-TAB and the BioStudies annotation model, respectively.
  • Researchers should consult funding agreements and target journals to determine specific repository mandates, as European Commission projects often require ENA/ArrayExpress, while NIH projects typically mandate SRA/GEO.

Researchers generating RNA-seq data face a practical decision early in their project: where to deposit raw sequencing reads, processed count matrices, and associated metadata. The choice affects how easily collaborators find your data, how smoothly journals accept your submission, and how efficiently you can retrieve comparable public datasets for reanalysis. This article compares the three major public repositories that host RNA-seq data, the National Center for Biotechnology Information Gene Expression Omnibus (NCBI GEO), the European Bioinformatics Institute ArrayExpress, and the European Nucleotide Archive (ENA), with attention to data coverage, submission requirements, search capabilities, and data formats.

The direct answer to which repository you should use depends on your primary need. For depositing raw sequencing reads, ENA is the designated archive for European-funded projects and accepts data from anywhere. For depositing processed expression matrices with rich sample annotations, GEO is the most widely used choice in the published literature. For retrieving curated functional genomics datasets, GEO and ArrayExpress both offer structured query interfaces, but GEO currently hosts a substantially larger collection of RNA-seq series. Many projects deposit raw reads in ENA and processed counts in GEO simultaneously, and this dual-deposit pattern is common in the literature.

Scope and Reader Context

This comparison serves biology students, researchers, laboratory professionals, and life-science practitioners who need to decide where to submit new RNA-seq data or where to search for existing datasets. The practical outcome is a decision framework based on your funding source, target journal, data type, and analysis workflow. The article covers the submission process for each repository, the search interfaces available for data retrieval, the file formats each repository accepts and distributes, and the metadata standards that govern discoverability.

The comparison does not cover every bioinformatics resource hosted by NCBI or EMBL-EBI. It focuses on the three repositories most relevant to RNA-seq workflows. Related resources such as the NCBI Sequence Read Archive (SRA) and the EMBL-EBI BioStudies database are discussed where they interact with the primary repositories.

At a Glance

FeatureNCBI GEOArrayExpressENA
Primary data typeProcessed expression matrices, microarray and RNA-seq seriesProcessed functional genomics data, microarray and RNA-seq seriesRaw sequencing reads, assembled sequences, and associated metadata
Submission routeGEO submission portal with SOFT, MINiML, or spreadsheet formatsBioStudies submission pipeline with ArrayExpress annotation spreadsheetsWebin submission system with ERA, SRA, or ENA file formats
Search interfaceGEO DataSets and GEO Profiles browsers with Entrez query systemBioconductor ArrayExpress package and EBI search portalENA browser with text search and programmatic access via ENA API
File formats acceptedProcessed data tables, raw data optional for microarray, raw sequence files routed to SRAProcessed data files, raw data routed to ENAFASTQ, BAM, CRAM, and other sequence formats
Metadata standardMINiML and SOFT formats with GEO sample and series attributesMAGE-TAB and BioStudies annotation modelINSDC standard with sample, experiment, run, and analysis objects
Typical use in RNA-seqDifferential expression analysis, meta-analysis, dataset discoveryFunctional genomics reanalysis, cross-species comparisonsRead alignment, variant calling, transcript assembly, data archiving
Accession prefixGSE for series, GSM for samples, GDS for datasetsE-MTAB for ArrayExpress experimentsERP for projects, ERX for experiments, ERR for runs
Funding mandate alignmentNIH and US-based grant requirementsEuropean Commission and UKRI requirementsEuropean Commission and global INSDC membership

Repository Roles in the RNA-seq Data Lifecycle

RNA-seq data exists in multiple forms throughout a project. Raw sequencing reads are produced by the instrument as FASTQ files. These reads are aligned to a reference genome or transcriptome, producing BAM files. Quantification tools generate count matrices that summarize expression levels for each gene or transcript. Downstream analysis produces differential expression tables, pathway enrichments, and visualizations.

The three repositories occupy different positions in this lifecycle. ENA archives raw sequencing reads and serves as the primary repository for sequence data in Europe. GEO archives processed expression data, including count matrices and clinical or experimental annotations, and is the dominant repository for functional genomics data in the United States. ArrayExpress historically served a similar role to GEO in Europe but has transitioned its submission pipeline to the BioStudies database while maintaining access to archived functional genomics datasets.

A researcher depositing RNA-seq data must decide whether to submit raw reads, processed counts, or both. Journals increasingly require raw sequencing data deposition as a condition of publication. Funding agencies often mandate deposition in specific repositories. The practical answer for most projects is to deposit raw reads in ENA or SRA and processed expression matrices in GEO or ArrayExpress.

NCBI GEO: The Dominant Repository for Processed Expression Data

The National Center for Biotechnology Information hosts the Gene Expression Omnibus as a public functional genomics data repository. GEO accepts array-based and high-throughput sequencing data related to gene expression, and it supports tools for querying and downloading experiments and curated gene expression profiles. The NCBI platform provides integrated search across GEO DataSets, GEO Profiles, and related sequence resources, which makes it a practical first stop for researchers seeking public RNA-seq data.

Data Coverage and Growth

GEO contains a large and growing collection of RNA-seq series. The repository has operated since the early 2000s and has accumulated submissions from diverse organisms, experimental designs, and disease contexts. Published studies frequently cite GEO accession numbers, and the repository serves as the default destination for processed expression data in the United States and internationally.

The volume of RNA-seq data in GEO reflects the broader adoption of the technology. Many studies deposit both raw sequence data in SRA and processed count matrices in GEO. This dual-deposit pattern means that a researcher searching GEO for a disease of interest will find series with sample annotations, experimental design details, and processed data tables that can be downloaded directly for reanalysis.

Submission Requirements and Process

GEO submission requires a structured set of metadata for each sample and series. The submission portal accepts data in several formats, including SOFT files, MINiML files, and Excel spreadsheets. For RNA-seq data, the processed count matrix is the primary data file, while raw sequence reads are routed to SRA as part of the submission process.

The GEO submission process requires the following components:

  • A series title and summary describing the experiment
  • Overall design information explaining the experimental groups and comparisons
  • Sample metadata including organism, tissue, cell type, treatment, and any relevant clinical variables
  • Processed data files containing expression measurements for each sample
  • Platform information describing the sequencing instrument and alignment approach

The submission system validates metadata fields and flags missing or inconsistent information. Submitters can update their records after deposition, which is useful when additional analyses are completed or when reviewers request additional metadata.

Search Capabilities and Data Retrieval

GEO offers multiple search interfaces. The GEO DataSets browser allows users to query by organism, platform, study type, and free-text terms. The GEO Profiles browser provides gene-level expression views across multiple datasets. The Entrez system integrates GEO with other NCBI databases, allowing users to move from a gene record to expression data for that gene across all GEO series.

For programmatic access, the NCBI provides the GEOquery package through Bioconductor. This package allows researchers to download GEO series directly into R for analysis. The Bioconductor project maintains documentation and workflows for using GEOquery with RNA-seq data, including steps for reading count matrices, accessing sample metadata, and preparing data for differential expression analysis.

Practical Use Cases for GEO

Researchers commonly use GEO for three purposes. First, they search for public datasets relevant to their biological question, download the processed count matrices, and perform their own differential expression analysis. Second, they use GEO series as validation cohorts to confirm findings from their own experiments. Third, they aggregate multiple GEO datasets for meta-analysis, combining expression data across studies to increase statistical power.

The published literature demonstrates these use patterns. Studies of bladder cancer have used gene and miRNA expression data from NCBI GEO to identify differentially expressed genes and build survival models. Research on Kawasaki disease has retrieved mRNA expression profiles from GEO to investigate the role of specific gene families. Studies of gastric mucosa after bariatric surgery have downloaded GEO series to identify differentially expressed genes and pathway enrichments.

ArrayExpress: The European Functional Genomics Archive

ArrayExpress is the functional genomics data repository at the European Bioinformatics Institute. It accepts submissions of microarray and high-throughput sequencing data and provides search and download services for processed expression data. ArrayExpress has historically served as the European counterpart to GEO, and many European journals and funding agencies have required or recommended deposition in this repository.

Current Status and Transition to BioStudies

ArrayExpress has undergone a significant transition in recent years. The submission pipeline for new functional genomics datasets has moved to the BioStudies database, which provides a broader framework for linking data files, protocols, and associated publications. Existing ArrayExpress datasets remain accessible, and the search interface continues to provide access to archived experiments.

This transition affects researchers planning new submissions. instead of submitting directly to ArrayExpress, new functional genomics datasets are submitted through BioStudies, with processed data files and metadata organized according to the BioStudies annotation model. The ArrayExpress accession format, such as E-MTAB, continues to be used for experiments archived in the system.

Data Coverage and Relationship to ENA

ArrayExpress contains a substantial collection of RNA-seq experiments, particularly from European research groups. The repository includes both microarray and high-throughput sequencing data, with processed expression matrices and sample annotations. For RNA-seq submissions, raw sequence reads are deposited in ENA, while processed data and experimental metadata are maintained in ArrayExpress or BioStudies.

The relationship between ArrayExpress and ENA mirrors the relationship between GEO and SRA. The processed data repository handles expression matrices and annotations, while the sequence archive handles raw reads. Researchers retrieving data from ArrayExpress receive processed count matrices, while researchers needing raw reads must access ENA.

Submission Requirements and Process

New submissions to ArrayExpress follow the BioStudies submission workflow. The process requires:

  • A study description with title, abstract, and experimental design
  • Sample annotations using controlled vocabulary where available
  • Processed data files in tabular or matrix format
  • Protocol descriptions for library preparation, sequencing, and data processing
  • Links to raw sequence data deposited in ENA

The submission system uses annotation spreadsheets to collect sample metadata. These spreadsheets follow a structured format with columns for sample characteristics, experimental factors, and technical variables. The system validates the spreadsheets and provides feedback on missing or malformed entries.

Search Capabilities and Data Retrieval

ArrayExpress provides a web search interface that supports queries by organism, experiment type, and free-text terms. The EBI search portal integrates ArrayExpress with other EMBL-EBI resources, allowing users to search across multiple databases from a single query.

For programmatic access, the Bioconductor ArrayExpress package provides tools for downloading experiments directly into R. This package supports searching for experiments by accession or keywords, downloading processed data files, and reading sample metadata into R data structures. The package documentation includes examples of retrieving RNA-seq experiments and preparing data for analysis.

Practical Use Cases for ArrayExpress

Researchers use ArrayExpress for similar purposes as GEO, including dataset discovery, validation, and meta-analysis. The repository is particularly relevant for researchers working with European collaborators or submitting to European journals. The EMBL-EBI training program offers courses and materials on using ArrayExpress and related resources for functional genomics analysis.

The published literature includes examples of ArrayExpress data use. Systems biology studies of cystic fibrosis have built gene regulatory networks from publicly available transcriptomic datasets, including those archived in European repositories. Studies of alternative splicing have used RNA-seq data from public repositories to characterize isoform diversity across organisms.

ENA: The European Nucleotide Archive for Raw Sequencing Reads

The European Nucleotide Archive is the European component of the International Nucleotide Sequence Database Collaboration, which also includes the NCBI Sequence Read Archive and the DNA Data Bank of Japan. ENA accepts raw sequencing reads, assembled sequences, and associated metadata, and it provides search and download services for sequence data.

Data Coverage and Role in the INSDC

ENA is the designated repository for raw sequencing data from European-funded projects. The archive accepts submissions from researchers worldwide and provides permanent accession numbers for sequence data. ENA is part of the INSDC, which means that data submitted to ENA is synchronized with SRA and DDBJ, allowing researchers to access the same data through any of the three archives.

For RNA-seq projects, ENA stores FASTQ files containing raw sequencing reads, along with metadata describing the sequencing instrument, library preparation, and sample characteristics. The archive also stores BAM and CRAM files for aligned reads, and it provides access to assembled transcript sequences when relevant.

Submission Requirements and Process

ENA submissions use the Webin submission system. The process requires:

  • A study or project registration describing the overall research project
  • Sample registrations with organism and sample attributes
  • Experiment registrations describing the sequencing library and instrument
  • Run registrations linking sequence files to experiments
  • Sequence file uploads in FASTQ, BAM, CRAM, or other accepted formats

The Webin system validates file formats and metadata fields, and it provides accession numbers at each stage of the submission. The system supports both interactive submissions through the web interface and programmatic submissions through the Webin REST API.

Search Capabilities and Data Retrieval

ENA provides a browser interface that supports text searches and filtering by organism, study type, and other attributes. The ENA API allows programmatic access to sequence data and metadata, supporting automated retrieval of large datasets.

For RNA-seq data, the ENA browser provides access to run files containing raw reads. Researchers can download FASTQ files for alignment and quantification, or they can access BAM files for aligned reads. The ENA API supports downloading files in bulk, which is useful for large-scale reanalysis projects.

Practical Use Cases for ENA

Researchers use ENA for three primary purposes. First, they deposit raw sequencing reads to satisfy journal and funding agency requirements. Second, they retrieve raw reads from public projects to perform their own alignment and quantification. Third, they access assembled sequences for transcript characterization.

The published literature demonstrates the importance of raw sequence data access. Studies of non-triplet alternative splicing have used RNA-seq data to characterize isoform diversity, requiring access to raw reads for alignment and quantification. Studies of cardiac fibrosis have conducted meta-analyses of public single-cell sequencing datasets, requiring retrieval of raw or processed data from multiple repositories.

Comparing Submission Workflows for RNA-seq Projects

The submission workflow for an RNA-seq project depends on which repository you choose and what data types you need to deposit. The following comparison covers the practical steps for each repository.

GEO Submission Workflow

Submitting processed RNA-seq data to GEO involves the following steps:

  1. Prepare processed count matrices for each sample, with genes as rows and samples as columns
  2. Prepare sample metadata, including organism, tissue, treatment, and experimental group
  3. Create a GEO submission using the web portal or spreadsheet templates
  4. Upload processed data files and metadata
  5. Submit raw sequence reads to SRA if required by the journal or funding agency
  6. Obtain GEO accession numbers for the series and samples
  7. Cite the GEO accession in your manuscript

The GEO submission system provides validation feedback during the submission process. Common issues include missing metadata fields, inconsistent sample names, and malformed data files. The system allows submitters to revise records after deposition.

ArrayExpress and BioStudies Submission Workflow

Submitting processed RNA-seq data through the BioStudies pipeline involves the following steps:

  1. Prepare processed count matrices and sample annotation spreadsheets
  2. Register the study in BioStudies and obtain a study accession
  3. Upload processed data files and annotation spreadsheets
  4. Submit raw sequence reads to ENA using the Webin system
  5. Link the ENA study accession to the BioStudies record
  6. Obtain the ArrayExpress or BioStudies accession for citation

The BioStudies submission system uses annotation spreadsheets that follow a structured format. The system validates the spreadsheets and provides feedback on missing or inconsistent entries. The submission process requires careful attention to controlled vocabulary terms for sample characteristics.

ENA Submission Workflow

Submitting raw sequencing reads to ENA involves the following steps:

  1. Register a study in the Webin system and obtain a project accession
  2. Register samples with organism and sample attributes
  3. Register experiments describing the sequencing library and instrument
  4. Upload sequence files in FASTQ, BAM, or CRAM format
  5. Obtain run accessions for each sequence file
  6. Link the ENA study accession to any processed data deposited in GEO or ArrayExpress

The Webin system provides validation of file formats and metadata. The system supports both interactive and programmatic submissions, and it provides accession numbers at each stage.

Search and Retrieval Comparison for RNA-seq Data

Finding existing RNA-seq data requires different search strategies for each repository. The following comparison covers the search interfaces and retrieval methods available.

Searching GEO for RNA-seq Data

The GEO DataSets browser supports queries by organism, study type, platform, and free-text terms. A typical search for RNA-seq data involves the following steps:

  1. Navigate to the GEO DataSets browser
  2. Enter search terms for the biological condition of interest
  3. Filter by organism and study type
  4. Review the list of matching series
  5. Open series records to review sample metadata and experimental design
  6. Download processed data files or use GEOquery to load data into R

The GEO Profiles browser provides gene-level views across datasets. This interface is useful for checking expression patterns of specific genes across multiple studies.

Searching ArrayExpress for RNA-seq Data

The ArrayExpress search interface supports queries by organism, experiment type, and free-text terms. The EBI search portal integrates ArrayExpress with other EMBL-EBI resources. A typical search involves the following steps:

  1. Navigate to the ArrayExpress search interface or the EBI search portal
  2. Enter search terms for the biological condition of interest
  3. Filter by organism and experiment type
  4. Review the list of matching experiments
  5. Open experiment records to review sample metadata and protocols
  6. Download processed data files or use the ArrayExpress Bioconductor package

Searching ENA for RNA-seq Data

The ENA browser supports text searches and filtering by organism, study type, and other attributes. The ENA API provides programmatic access. A typical search involves the following steps:

  1. Navigate to the ENA browser
  2. Enter search terms for the biological condition of interest
  3. Filter by organism and data type
  4. Review the list of matching studies and runs
  5. Open run records to review sequencing metadata
  6. Download FASTQ files or use the ENA API for bulk retrieval

Data Formats and Compatibility with Analysis Workflows

The format of data retrieved from each repository affects how easily it can be used in downstream analysis workflows. The following comparison covers the formats available from each repository.

Processed Data Formats

GEO provides processed data in SOFT and MINiML formats, along with supplementary files that often contain count matrices in tabular format. The GEOquery Bioconductor package reads these formats and converts them into R data structures suitable for analysis.

ArrayExpress provides processed data in MAGE-TAB format, along with supplementary files containing expression matrices. The ArrayExpress Bioconductor package reads these formats and loads data into R.

ENA provides raw sequence data in FASTQ, BAM, and CRAM formats. These formats require alignment and quantification before expression analysis can proceed.

Raw Sequence Data Formats

ENA and SRA store raw sequencing reads in FASTQ format, along with aligned reads in BAM or CRAM format. These files are typically large, ranging from hundreds of megabytes to hundreds of gigabytes per sample.

The choice between processed and raw data affects the analysis workflow. Processed count matrices can be used directly for differential expression analysis, while raw reads require alignment and quantification steps. The Galaxy Training Network provides tutorials for RNA-seq analysis workflows that start from raw reads and produce count matrices. The nf-core project provides community-developed pipelines for RNA-seq analysis that accept raw reads as input and produce count matrices and quality reports.

Metadata Formats and Standards

The metadata standards used by each repository affect data discoverability and interoperability. GEO uses the MINiML and SOFT formats, which encode sample attributes, experimental design, and data processing information. ArrayExpress uses the MAGE-TAB format, which provides a structured representation of experimental metadata. ENA uses the INSDC standard, which defines sample, experiment, run, and analysis objects.

The FAIR guiding principles for data management and reuse emphasize that data should be findable, accessible, interoperable, and reusable. The metadata standards used by these repositories support these principles by providing structured descriptions of experimental context. Studies that reuse public omics data have demonstrated the value of well-annotated repositories for enabling new research findings.

Practical Implementation Steps for Repository Selection

Selecting the appropriate repository for your RNA-seq project requires a systematic assessment of your requirements. The following steps provide a practical framework for this decision.

Step 1: Review Funding and Journal Requirements

Check your funding agreement and target journal for data deposition requirements. European Commission funded projects typically require deposition in ENA for raw reads and ArrayExpress or BioStudies for processed data. NIH funded projects typically require deposition in SRA for raw reads and GEO for processed data. Many journals require raw data deposition as a condition of publication.

Step 2: Determine Your Data Types

Identify whether your project generates raw sequencing reads, processed count matrices, or both. Raw reads require deposition in ENA or SRA. Processed count matrices require deposition in GEO or ArrayExpress. Most RNA-seq projects generate both data types and require dual deposition.

Step 3: Assess Metadata Requirements

Review the metadata requirements for each repository. GEO requires detailed sample annotations and experimental design information. ArrayExpress and BioStudies require annotation spreadsheets with controlled vocabulary terms. ENA requires sample attributes and sequencing metadata. Choose the repository where you can provide the required metadata without excessive burden.

Step 4: Evaluate Search and Retrieval Needs

Consider whether you need to retrieve data from the repository for your own analysis. GEO provides the most mature search interface for processed expression data. ENA provides the most comprehensive access to raw sequence data. ArrayExpress provides access to European functional genomics datasets.

Step 5: Plan for Dual Deposition

For most RNA-seq projects, plan to deposit raw reads in ENA or SRA and processed data in GEO or ArrayExpress. This dual-deposit pattern satisfies journal and funding requirements while maximizing data discoverability. Link the raw and processed data records through accession numbers.

Records and Measurements for Repository Selection

Maintaining records of your repository selection and submission process supports reproducibility and compliance. The following records are recommended for RNA-seq projects.

Submission Records

Document the following information for each repository submission:

  • Repository name and accession numbers
  • Submission date and submission system used
  • Data files uploaded and their formats
  • Metadata fields provided and any controlled vocabulary terms used
  • Validation feedback received and corrections made
  • Links between raw and processed data records

Data Retrieval Records

Document the following information for each data retrieval:

  • Repository name and accession numbers of retrieved datasets
  • Search terms and filters used
  • Download date and file formats retrieved
  • Software versions used for data processing
  • Any data quality issues observed

Analysis Records

Document the following information for each analysis using public data:

  • Source repository and accession numbers
  • Data processing steps and software versions
  • Quality control metrics and filtering decisions
  • Analysis parameters and statistical methods
  • Output files and their locations

Common Failure Patterns in Repository Selection and Use

Researchers encounter several common problems when selecting and using RNA-seq repositories. Recognizing these patterns helps avoid costly delays and data loss.

Failure Pattern 1: Submitting Only Processed Data

Some researchers deposit only processed count matrices in GEO or ArrayExpress and skip raw read deposition. This pattern fails to satisfy journal and funding requirements that mandate raw data availability. Reviewers may request raw reads for independent verification of results. The solution is to deposit raw reads in ENA or SRA alongside processed data.

Failure Pattern 2: Submitting Only Raw Reads

The opposite pattern involves depositing only raw reads in ENA or SRA without processed count matrices. This pattern makes it difficult for other researchers to reuse the data, because they must perform their own alignment and quantification before they can interpret the results. The solution is to deposit processed count matrices in GEO or ArrayExpress with appropriate sample annotations.

Failure Pattern 3: Incomplete Metadata

Submissions with incomplete or inconsistent metadata are difficult to discover and reuse. Missing sample attributes, unclear experimental design descriptions, and inconsistent sample names all reduce data utility. The solution is to follow the metadata standards for each repository and to provide comprehensive annotations for every sample.

Failure Pattern 4: Ignoring Repository-Specific Requirements

Each repository has specific submission requirements and file format expectations. Submissions that ignore these requirements are rejected or delayed. The solution is to review the submission documentation for each repository before starting the submission process.

Failure Pattern 5: Failing to Link Raw and Processed Data

Dual depositions that do not link raw and processed data records create confusion for data reusers. Researchers may find the processed data but not know where to access the raw reads, or they may find the raw reads but not know where to access the processed counts. The solution is to include cross-references between repository records.

Quality Controls for Repository Data

Quality control is essential when using public RNA-seq data from any repository. The following controls help ensure that retrieved data is suitable for analysis.

Metadata Quality Assessment

Review the metadata for each dataset before downloading data. Check that sample annotations are complete and consistent with the experimental design. Look for missing values, inconsistent terminology, and unclear descriptions. Datasets with poor metadata may not be suitable for your analysis.

Data Quality Assessment

Assess the quality of processed count matrices before analysis. Check for missing values, zero counts, and unusual distributions. Compare sample-level metrics such as total read counts and gene detection rates. The SkewC tool provides a method for identifying cells with skewed gene body coverage in single-cell RNA-seq data, which is relevant for quality assessment in scRNA-seq workflows.

Reproducibility Assessment

Check whether the dataset includes sufficient information to reproduce the analysis. Look for descriptions of alignment methods, quantification tools, and filtering criteria. Datasets that lack processing details may not be suitable for meta-analysis or validation studies.

Limitations of Repository Data

Public RNA-seq data from any repository has limitations that affect its utility for reanalysis. Understanding these limitations helps researchers interpret results appropriately.

Batch Effects and Technical Variation

RNA-seq data generated in different laboratories, at different times, and with different protocols exhibit technical variation that can confound biological comparisons. Meta-analyses that combine data from multiple studies must account for batch effects through statistical methods or study design.

Incomplete Metadata

Many public datasets lack complete metadata, making it difficult to interpret experimental conditions or to identify appropriate samples for comparison. Researchers must carefully review available metadata and document any assumptions made during data selection.

Processing Pipeline Differences

Different studies use different alignment and quantification pipelines, producing count matrices that are not directly comparable. Researchers must decide whether to use processed data as provided or to reprocess raw reads with a consistent pipeline.

Sample Size Limitations

Public datasets often have small sample sizes, particularly for rare diseases or specialized experimental conditions. Small sample sizes limit statistical power and increase the risk of false discoveries.

Safety and Regulatory Context

Data deposition in public repositories has regulatory and ethical implications that researchers must consider.

Data Privacy and Consent

RNA-seq data derived from human samples must comply with privacy regulations and consent requirements. Researchers must ensure that deposited data does not contain identifiable information and that the original consent covers public data sharing. Some repositories provide controlled access options for sensitive data.

Data Use Agreements

Some datasets are subject to data use agreements that restrict how the data can be used. Researchers retrieving data from public repositories must review and comply with any applicable data use agreements.

Export Controls

Raw sequencing data may be subject to export controls in some jurisdictions. Researchers depositing or retrieving data across international borders must comply with applicable regulations.

Professional Escalation Criteria

Researchers should seek professional guidance when encountering specific situations during repository selection or data use.

When to Consult a Bioinformatics Specialist

Consult a bioinformatics specialist when you need to:

  • Choose between processed and raw data for your analysis
  • Design a reprocessing pipeline for raw reads from public datasets
  • Integrate data from multiple repositories with different formats
  • Address batch effects in meta-analyses of public data

When to Consult a Data Manager or Librarian

Consult a data manager or librarian when you need to:

  • Understand funding agency data deposition requirements
  • Prepare metadata for repository submission
  • Navigate controlled access data request processes
  • Comply with data privacy regulations

When to Consult an Ethics or Regulatory Advisor

Consult an ethics or regulatory advisor when you need to:

  • Determine whether your data deposition plan complies with consent requirements
  • Assess whether your data contains identifiable information
  • Navigate international data transfer regulations
  • Address data use agreement restrictions

Frequently Asked Questions

What is the main difference between GEO and ENA for RNA-seq data?

GEO stores processed expression data, including count matrices and sample annotations, while ENA stores raw sequencing reads in FASTQ, BAM, or CRAM format. For a complete RNA-seq submission, you typically deposit raw reads in ENA and processed counts in GEO. The two repositories serve complementary roles in the data lifecycle.

Can I submit the same RNA-seq dataset to both GEO and ArrayExpress?

You can deposit processed data in both GEO and ArrayExpress, but this is rarely necessary. Most projects deposit processed data in one repository and raw reads in the corresponding sequence archive. The choice between GEO and ArrayExpress typically depends on your funding source and target journal.

Do I need to deposit raw sequencing reads if I already deposited processed counts in GEO?

Most journals and funding agencies require raw sequencing read deposition in addition to processed counts. Raw reads allow other researchers to independently verify your analysis and to apply alternative processing methods. Check your target journal and funding agreement for specific requirements.

How do I search for RNA-seq datasets in GEO?

Use the GEO DataSets browser to search by organism, study type, and free-text terms. Filter results by platform and study design. Open series records to review sample metadata and download processed data files. The GEOquery Bioconductor package provides programmatic access for downloading data into R.

What file formats does ENA accept for RNA-seq data?

ENA accepts FASTQ files for raw sequencing reads, BAM and CRAM files for aligned reads, and FASTA files for assembled sequences. The Webin submission system validates file formats and provides accession numbers for each submission stage.

How long does it take to submit RNA-seq data to these repositories?

Submission time varies depending on data size and metadata complexity. Processed data submissions to GEO or ArrayExpress can be completed in hours if metadata is prepared in advance. Raw read submissions to ENA can take longer due to file upload times for large datasets.

Can I access data from these repositories programmatically?

Yes, all three repositories support programmatic access. GEO provides the GEOquery Bioconductor package and the Entrez API. ArrayExpress provides the ArrayExpress Bioconductor package and the EBI search API. ENA provides the ENA API for bulk data retrieval.

What should I do if I find errors in a public dataset?

Contact the data submitter through the repository contact information to report errors. Many submitters appreciate feedback and may update their records. If the errors affect your analysis, document the issues and consider excluding the dataset from your analysis.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.