Zubair Khalid

Virologist/Molecular Biologist | Veterinarian | Bioinformatician

Conventional & Molecular Virology • Vaccine Development • Computational Biology

Dr. Zubair Khalid is a veterinarian and virologist specializing in conventional and molecular virology, vaccine development, and computational biology. Dedicated to advancing animal health through innovative research and multi-omics approaches.

Dr. Zubair Khalid - Veterinarian, Virologist, and Vaccine Development Researcher specializing in Computational Biology, Multi-omics, Animal Health, and Infectious Disease Research

Section: Infrastructure, Cloud & Policy

RNA-Seq Databases: Accessing and Using Public RNA-Seq Data

Public RNA-seq databases hold millions of transcriptomic profiles generated by research groups worldwide, and these datasets are available for reuse in meta-analyses, validation studies, and new hypothesis testing. This article explains how students, researchers, analysts, and life-science professionals can locate, download, and repurpose public RNA-seq data through major repositories including the Gene Expression Omnibus (GEO), the Sequence Read Archive (SRA), ENCODE, and the EMBL-EBI Expression Atlas. The practical outcome is a working workflow for searching these databases, evaluating dataset quality, downloading raw or processed data, and importing it into common analysis tools.

The Scale and Value of Public RNA-Seq Repositories

Advances in next generation sequencing technologies have produced a broad array of large-scale gene expression studies and an unprecedented volume of whole messenger RNA sequencing data, also known as RNA-seq [5]. Major projects such as the Genotype Tissue Expression project (GTEx) and The Cancer Genome Atlas (TCGA) have deposited thousands of samples into public archives [5]. Beyond these flagship projects, individual studies contribute tens of thousands of additional samples across species, tissues, and experimental conditions.

The volume of deposited data creates both opportunity and challenge. Millions of transcriptomic profiles have been placed in public archives, yet many remain underused for the interpretation of new experiments [12]. Researchers who learn to navigate these repositories gain access to validation cohorts, control datasets, and cross-species comparisons that would be prohibitively expensive to generate independently.

Public RNA-seq data supports several practical applications. Researchers can validate their own differential expression findings against independent cohorts. Analysts can perform meta-analyses across multiple studies to increase statistical power. Computational groups can build reference transcriptomes from aggregated public data. For example, one group constructed reference transcriptome data for the Western honeybee using the reference genome sequence and RNA-seq data curated from about 1,000 runs of public databases, producing 149,685 transcripts and 194,174 predicted protein sequences [11]. This approach demonstrates how public data can be repurposed to create resources that benefit entire research communities.

Major Public RNA-Seq Databases

Gene Expression Omnibus (GEO)

The Gene Expression Omnibus is a public functional genomics data repository maintained by the National Center for Biotechnology Information (NCBI) [2]. GEO accepts array-based and high-throughput sequencing data, including RNA-seq. Each submission receives a series accession number such as GSE89408, which researchers can cite in publications and use to retrieve the complete dataset [9].

GEO organizes data into three levels. A series (GSE) represents a complete study. A sample (GSM) represents one biological or technical replicate. A platform (GPL) describes the technology used. For RNA-seq studies, the platform entry typically describes the sequencing instrument and library preparation method.

Searching GEO effectively requires understanding its query syntax. The GEO DataSets interface supports Boolean operators, field tags such as "organism"[Organism] and "dataset type"[DataSet Type], and the ability to filter by entry type, organism, and study type. Users can also browse GEO by organism and platform.

Sequence Read Archive (SRA)

The Sequence Read Archive is NCBI's repository for raw sequencing reads [2]. While GEO stores processed data and metadata, SRA stores the primary sequence data generated by sequencing instruments. For RNA-seq studies, SRA files contain the actual FASTQ reads that must be aligned to a reference genome or transcriptome before expression quantification.

SRA data can be accessed through the SRA Run Selector, which allows users to browse samples within a study, view metadata, and download files. The SRA Toolkit provides command-line utilities including fasterq-dump and prefetch for downloading and converting SRA files to FASTQ format.

The relationship between GEO and SRA is important for practical use. Many RNA-seq studies deposit processed count matrices in GEO and raw reads in SRA. Researchers who need to reanalyze data with a different alignment or quantification method must download raw reads from SRA. Researchers who only need gene-level counts can often use the processed data from GEO.

ENCODE

The Encyclopedia of DNA Elements (ENCODE) is a public research project that aims to identify functional elements in the human and mouse genomes. ENCODE has generated extensive RNA-seq data across many cell types and experimental conditions. The ENCODE portal provides access to uniformly processed data, including gene expression quantifications and raw reads.

ENCODE data is particularly valuable for researchers studying gene regulation because it includes matched assays such as chromatin accessibility and transcription factor binding. The uniform processing pipelines used by ENCODE reduce the batch effects that complicate cross-study comparisons.

EMBL-EBI Expression Atlas

The Expression Atlas is a database maintained by the European Bioinformatics Institute (EMBL-EBI) that provides information about gene expression patterns across species, tissues, cell types, and experimental conditions [1]. The Expression Atlas includes both baseline expression data and differential expression results from curated studies.

The Expression Atlas differs from GEO and SRA in its focus on curated, reanalyzed data. EMBL-EBI curators reprocess raw data through standardized pipelines, which makes cross-study comparisons more reliable. The database supports queries by gene, condition, or biological process, and results are displayed as expression heatmaps and bar charts.

The European Bioinformatics Institute also provides training materials for using these resources effectively [1]. Researchers who are new to public data analysis should review these materials before beginning large-scale downloads.

Species-Specific and Domain-Specific Databases

Beyond the major general-purpose repositories, several specialized databases aggregate RNA-seq data for particular species or research questions. These resources often provide additional processing and visualization features.

For livestock research, a web-based database server using 43,710 public RNA-seq samples supports analysis of gene expression and alternative splicing in livestock animals [19]. This resource allows researchers to query expression levels across tissues and breeds without downloading and processing raw data themselves.

For plant research, the Plant Intron-Splicing Efficiency Database (PISE) explores splicing of approximately 1,650,000 introns in Arabidopsis, maize, rice, and soybean from about 57,000 public RNA-seq libraries [20]. This database provides precomputed splicing efficiency values that would be computationally expensive to generate independently.

For grape research, Grape-rna is a database for the collection, evaluation, treatment, and data sharing of grape RNA-seq datasets [21]. This resource addresses the need for species-specific data curation and quality control.

Domain-specific databases also exist for disease-focused research. BRPtools is a web platform that manually curated and uniformly reprocessed 134 publicly available human RNA-seq datasets covering 88 distinct blood-related diseases from whole blood and peripheral blood mononuclear cell data [14]. The platform includes a search module for exploring gene and disease level expression profiles and a predict module for disease classification using user-uploaded RNA-seq count matrices [14].

At a Glance: Database Selection Guide

Database Data Type Access Method Best Use Case Processing Level
GEO Processed counts, metadata Web search, GEOquery R package Differential expression validation, metadata mining Processed data readily available
SRA Raw sequencing reads SRA Toolkit, fasterq-dump Reanalysis with custom pipelines, alignment to new references Raw FASTQ files
ENCODE Uniformly processed RNA-seq, matched assays ENCODE portal, file downloads Gene regulation studies, cell type comparisons Uniform pipelines
Expression Atlas Curated baseline and differential expression Web queries, API access Cross-species and cross-tissue comparisons Reanalyzed by curators
Species-specific databases Processed expression, splicing Web interfaces Livestock, plant, or crop specific queries Precomputed values

Searching Public RNA-Seq Databases

Developing a Search Strategy

Effective searching begins with a clear biological question. Before opening a database interface, define the organism, tissue or cell type, condition or treatment, and the type of data needed. This clarity prevents wasted time browsing unrelated studies.

For GEO searches, combine organism and study type filters with keyword terms. A search for "muscle AND exercise AND Homo sapiens"[Organism] would return human studies examining exercise effects on muscle tissue. The GEO DataSets interface displays search results with study titles, summaries, sample counts, and submission dates.

For SRA searches, use the SRA Run Selector to filter by organism, platform, and study attributes. The SRA database also supports text searches for experimental design terms. Users can save search results and download metadata tables for offline review.

The Expression Atlas supports gene-centric queries. Users can search for a gene symbol and view its expression across all curated datasets. This approach is useful for checking whether a gene of interest is expressed in a particular tissue or condition before designing experiments.

Evaluating Dataset Relevance

Not all search results will be suitable for a given analysis. Evaluate each candidate dataset by reviewing the study summary, experimental design, sample annotations, and data processing methods. Key questions include:

Does the study include appropriate control samples? Are the sample sizes adequate for the intended analysis? Were the samples collected from the same tissue or cell type as the research question? What library preparation and sequencing platform were used?

For meta-analyses, document the inclusion and exclusion criteria before searching. This documentation should specify the minimum sample size, the required metadata fields, and the acceptable range of sequencing platforms. Predefined criteria reduce the risk of selection bias.

Using Metadata to Filter Results

Metadata quality varies substantially across public datasets. Some studies provide detailed sample annotations including age, sex, treatment dose, and time point. Other studies provide minimal information that makes interpretation difficult.

The FAIR Guiding Principles describe the importance of making data findable, accessible, interoperable, and reusable [4]. Datasets that follow these principles include rich metadata, use standard file formats, and provide clear documentation of processing steps. When comparing candidate datasets, prefer those with complete metadata because they support more robust downstream analysis.

Downloading RNA-Seq Data

Processed Data from GEO

Processed RNA-seq data from GEO is typically available as supplementary files attached to the series record. These files may include gene-level count matrices, normalized expression values, or transcript-level quantifications. The format depends on the submitting group and the analysis pipeline used.

The GEOquery R package provides programmatic access to GEO data. This package can download series matrix files, parse sample metadata, and retrieve supplementary files. For researchers who prefer command-line tools, the NCBI e-utilities API supports automated downloads.

Raw Reads from SRA

Raw RNA-seq reads are stored in SRA format, which is a compressed representation of the original FASTQ files. The SRA Toolkit provides several download options:

The prefetch command downloads SRA files to a local cache. The fasterq-dump command converts SRA files to FASTQ format. The sam-dump command outputs aligned reads in SAM format for studies that include alignments.

Downloading raw reads requires substantial storage space. A single RNA-seq sample can require several gigabytes of FASTQ data. For large studies with hundreds of samples, plan for terabytes of storage. Check the SRA Run Selector for file sizes before beginning a download.

Uniformly Processed Data from ENCODE and Expression Atlas

ENCODE provides processed data files that can be downloaded directly from the portal. These files include gene quantifications in TSV format and raw FASTQ files for users who need to rerun alignments.

The Expression Atlas provides processed expression data through its web interface and API. Users can download expression matrices for individual studies or across multiple studies. The API supports programmatic queries for large-scale analyses.

Importing Public RNA-Seq Data into Analysis Tools

Preparing Count Matrices

Most downstream analyses begin with a count matrix where rows represent genes or transcripts and columns represent samples. Public databases provide count matrices in various formats, and these formats often require conversion before use in analysis tools.

For GEO data, the supplementary files may contain raw counts, normalized counts, or both. Raw counts are required for differential expression analysis with tools such as DESeq2 or edgeR. Normalized counts are appropriate for visualization and exploratory analysis.

For SRA data, the raw reads must be aligned and quantified before count matrices can be generated. This process involves quality trimming, alignment to a reference genome or transcriptome, and quantification. Common tools include STAR, HISAT2, and Salmon. The choice of alignment and quantification method affects the resulting count values, so document these choices carefully.

Handling Batch Effects

Public RNA-seq data from different studies often show systematic technical variation known as batch effects. These effects arise from differences in library preparation protocols, sequencing instruments, and sample processing dates. Batch effects can obscure biological signals or create false associations if not addressed.

Several strategies mitigate batch effects in meta-analyses. Including study or batch as a covariate in statistical models accounts for systematic differences. ComBat and other empirical Bayes methods can remove batch effects before downstream analysis. The GenomicSuperSignature method applies Principal Component Analysis across 536 studies comprising 44,890 human RNA sequencing profiles and aggregates sufficiently similar loading vectors to form Replicable Axes of Variation [12]. This approach enables comparison of new datasets to public databases while remaining robust to batch effects and heterogeneous training data [12].

Using R and Bioconductor

The R programming language and Bioconductor packages provide the most common environment for analyzing public RNA-seq data. Key packages include:

DESeq2 for differential expression analysis of count data. edgeR for differential expression analysis with empirical Bayes methods. limma for linear model based analysis of microarray and RNA-seq data. clusterProfiler for functional enrichment analysis. Seurat for single-cell RNA-seq analysis.

The GenomicSuperSignature R/Bioconductor package demonstrates the value of integrating public data comparison into standard analysis workflows [12]. This package associates new datasets with annotated axes of variation, extracts interpretable annotations, and provides intuitive visualization [12].

Using Command-Line Tools

For large-scale analyses, command-line tools offer efficiency and reproducibility. The SRA Toolkit, STAR aligner, HISAT2 aligner, and Salmon quantifier can be combined into processing pipelines. Workflow managers such as Snakemake and Nextflow help organize these pipelines and track processing steps.

A typical command-line workflow for raw SRA data includes:

Download SRA files with prefetch. Convert to FASTQ with fasterq-dump. Trim adapters and low-quality bases with fastp or Trimmomatic. Align reads to the reference genome with STAR or HISAT2. Quantify gene or transcript expression with featureCounts or Salmon. Generate a count matrix for downstream analysis.

Practical Workflow for Reusing Public RNA-Seq Data

Step 1: Define the Analysis Question

Write a clear statement of the biological question and the type of data needed to answer it. Specify the organism, tissue, condition, and the required sample size. Determine whether processed counts or raw reads are necessary.

Step 2: Search Candidate Databases

Search GEO, SRA, ENCODE, and the Expression Atlas using the search strategies described above. Record the search terms, filters, and dates for reproducibility. Save the list of candidate studies with their accession numbers.

Step 3: Screen Studies for Eligibility

Review each candidate study against the predefined inclusion and exclusion criteria. Document the reason for excluding any study. For included studies, record the sample size, experimental design, and data processing methods.

Step 4: Download Data

Download processed count matrices from GEO or the Expression Atlas when available. Download raw reads from SRA when reanalysis is required. Verify file integrity by checking file sizes and checksums.

Step 5: Process and Analyze

Import count matrices into the chosen analysis tools. Apply quality control filters to remove low-count genes and low-quality samples. Perform differential expression, clustering, or other analyses according to the research question.

Step 6: Validate and Interpret

Compare results across studies to identify consistent findings. Use the GenomicSuperSignature approach or similar methods to place new results in the context of public data [12]. Interpret findings with attention to the limitations of the source data.

Step 7: Document and Share

Record all accession numbers, processing steps, and software versions. Share the analysis code and processed data to support reproducibility. Cite the source databases and original studies in publications.

Records and Measurements for Reproducible Analysis

Maintaining an Analysis Log

Reproducible analysis of public RNA-seq data requires careful record keeping. Maintain an analysis log that includes:

The date of each search and the search terms used. The accession numbers of all downloaded datasets. The software versions used for processing and analysis. The parameters used for alignment, quantification, and statistical analysis. The file paths for all input and output files.

This log serves as the primary record for reproducing the analysis and for responding to reviewer questions.

Tracking Data Versions

Public databases occasionally update records. A study that was downloaded in one version may differ from the same study downloaded later. Record the download date and the database version for each dataset. Check for updates before publishing results.

Measuring Data Quality

Quality metrics for RNA-seq data include:

The number of raw reads per sample. The percentage of reads that align to the reference genome. The percentage of reads that map to exonic regions. The number of genes detected at a minimum expression threshold. The distribution of expression values across samples.

These metrics help identify low-quality samples that should be excluded from analysis. Compare quality metrics across samples within a study and across studies in a meta-analysis.

Common Failure Patterns in Public Data Reuse

Insufficient Metadata

Many public datasets lack detailed sample annotations. Without information about treatment conditions, sample collection methods, or data processing steps, the data cannot be interpreted correctly. This limitation should be documented and considered when drawing conclusions.

Batch Effects Confounded with Biological Signals

When samples from different studies are combined, technical variation can correlate with the biological variable of interest. For example, if all control samples come from one study and all treated samples from another, the analysis cannot distinguish treatment effects from batch effects. This confounding is a common cause of false findings in meta-analyses.

Misaligned Reference Genomes

Different studies may align reads to different versions of the reference genome or transcriptome. Gene identifiers and coordinates can differ between versions, making cross-study comparisons difficult. Convert all data to a common reference version before analysis.

Inconsistent Gene Annotations

Public datasets use various gene annotation sources including RefSeq, Ensembl, and GENCODE. Gene symbols and identifiers may not match across sources. Use a consistent annotation for all samples and convert identifiers when necessary.

Overlooking Data Use Restrictions

Some datasets have restrictions on use beyond the standard NIH Genomic Data Sharing Policy [3]. Review the data use agreements for each dataset before beginning analysis. Document any restrictions in the analysis log.

Limitations of Public RNA-Seq Data

Technical Variability

Public RNA-seq data comes from many laboratories using different protocols. Library preparation methods, sequencing platforms, read lengths, and depth vary across studies. These technical differences contribute to batch effects and reduce the comparability of samples from different studies.

Biological Heterogeneity

Samples within a study may come from heterogeneous patient populations or animal cohorts. Genetic background, age, sex, and environmental factors influence gene expression. Without detailed metadata, these sources of variation cannot be fully accounted for in analysis.

Selection Bias

Public databases contain a nonrandom sample of all possible experiments. Studies with positive results are more likely to be published and deposited. This publication bias can affect the conclusions drawn from meta-analyses of public data.

Annotation Completeness

Reference genome and transcriptome annotations are incomplete for many species. Novel transcripts and isoforms may be missed by alignment based quantification methods. The honeybee reference transcriptome construction identified 149,685 transcripts and predicted 194,174 protein sequences, with approximately 50 to 60 percent of the predicted protein sequences functionally annotated using protein sequence data from several model and insect species [11]. This example shows that even well-annotated species have substantial uncharacterized transcript diversity.

Quality and Welfare Context for Animal Studies

Ethical Use of Animal-Derived Data

Public RNA-seq data from animal studies represents a valuable resource that can reduce the need for new animal experiments. Reusing existing data aligns with the principles of replacement, reduction, and refinement in animal research. Researchers should consider whether public data can answer their question before proposing new animal studies.

Data from Livestock and Agricultural Species

Public RNA-seq data for livestock species supports research on production traits, disease resistance, and product quality. A study of unsaturated fatty acid generation in buffaloes conducted RNA-seq on adipose tissue samples from six distinct depots and identified 8,926 lncRNAs, of which 1,363 were novel [13]. The study revealed a module including 207 lncRNAs with high correlation to unsaturated fatty acid content and identified two lncRNAs with predominant expression in sternum subcutaneous adipose tissue [13]. These findings demonstrate how public data reuse can generate insights relevant to beef quality and nutritional value.

Regulatory Considerations

Researchers who generate new RNA-seq data from human subjects must comply with the NIH Genomic Data Sharing Policy [3]. This policy establishes expectations for data sharing, privacy protection, and responsible use of genomic data. Researchers who reuse public data must also comply with the terms of the original data use agreements.

The FAIR Guiding Principles provide a framework for making data findable, accessible, interoperable, and reusable [4]. Researchers who deposit new data should follow these principles to maximize the value of their contributions to the scientific community.

Professional Escalation Criteria

When to Seek Expert Assistance

Several situations warrant consultation with a bioinformatics specialist or biostatistician:

When the analysis requires integration of data from more than five studies with different platforms or protocols. When batch effects are severe enough to obscure biological signals. When the research question requires custom analysis methods not available in standard tools. When the results will inform regulatory submissions or clinical decisions.

When to Reconsider the Analysis Approach

Reconsider the approach when:

The available public data does not include appropriate control samples. The sample size is too small to detect the expected effect sizes. The metadata is insufficient to support the intended comparisons. The technical variability across studies exceeds the biological signal of interest.

When to Generate New Data

Generating new RNA-seq data may be necessary when:

The biological question requires specific experimental conditions not represented in public data. The available public data lacks the necessary sample types or time points. The quality of existing data is too poor for reliable analysis. The research requires matched samples from the same individuals or animals.

Frequently Asked Questions

What is the difference between GEO and SRA?

GEO stores processed data and study metadata, while SRA stores raw sequencing reads. For RNA-seq studies, GEO typically contains gene-level count matrices and normalized expression values. SRA contains the original FASTQ reads that must be aligned and quantified before analysis. Researchers who need to reanalyze data with a different pipeline must download from SRA, while researchers who only need expression values can use GEO.

How do I find RNA-seq datasets for a specific tissue or condition?

Use the search interfaces of GEO, SRA, and the Expression Atlas with organism, tissue, and condition terms. The Expression Atlas supports gene-centric queries that show expression across all curated datasets. For species-specific resources, consult databases such as the livestock RNA-seq database or PISE for plants [19][20].

Can I combine RNA-seq data from different studies in a meta-analysis?

Yes, but combining data from different studies requires careful handling of batch effects. Include study as a covariate in statistical models or use batch correction methods such as ComBat. The GenomicSuperSignature method provides an alternative approach that is robust to batch effects and heterogeneous training data [12].

What are the minimum metadata requirements for reusing public RNA-seq data?

The minimum metadata includes the organism, tissue or cell type, experimental condition, and sample collection method. The FAIR Guiding Principles describe the importance of rich metadata for data reuse [4]. Datasets with incomplete metadata may still be usable but require careful interpretation.

How much storage space do I need for raw RNA-seq data?

A single RNA-seq sample typically requires several gigabytes of FASTQ data. A study with 100 samples may require hundreds of gigabytes. Check the SRA Run Selector for file sizes before downloading. Processed count matrices require much less storage, typically a few megabytes per study.

What quality metrics should I check before using public RNA-seq data?

Check the number of raw reads per sample, the alignment rate to the reference genome, the percentage of reads mapping to exonic regions, and the number of genes detected. Compare these metrics across samples to identify outliers. Document all quality metrics in the analysis log.

Are there restrictions on using public RNA-seq data?

Some datasets have data use restrictions beyond the standard NIH Genomic Data Sharing Policy [3]. Review the data use agreements for each dataset before analysis. Some datasets require approval from the original investigators or restrict use to specific research areas.

How do I cite public RNA-seq data in my publications?

Cite the original study that generated the data using its publication reference and the database accession number. For GEO data, cite the GSE accession number. For SRA data, cite the SRA study accession. Include the database name and accession in the methods section of the publication.

Related Bioinformatics Guides

References and Further Reading

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.