Single-Cell Sequencing Databases: Resources for Data Sharing and Exploration
Single-cell sequencing has transformed biological research by enabling transcriptome analysis at the resolution of individual cells. The first single-cell mRNA sequencing assays demonstrated that a single mouse blastomere could yield detection of thousands more genes than microarray techniques, revealing transcript isoform complexity that bulk methods could not resolve [6]. As these technologies have matured, the volume of publicly available single-cell data has grown exponentially, creating both opportunity and challenge for researchers who need to locate, access, and reuse datasets. This article provides a practical orientation to the major databases and repositories for single-cell sequencing data, with attention to access policies, data formats, and the practical decisions researchers face when depositing or retrieving data.
The Data Landscape in Single-Cell Genomics
Single-cell RNA sequencing (scRNA-seq) has become a standard tool for characterizing cell states and heterogeneity, but the analytical value of any single experiment depends on the ability to compare findings against existing data [5]. Public repositories serve as the primary infrastructure for this comparison, hosting raw sequencing files, processed count matrices, and increasingly, cell-type annotations and spatial transcriptomics data.
The scale of available data presents both opportunities and complications. Large atlas projects aggregate samples across laboratories, locations, and experimental conditions, introducing complex batch effects that require careful integration methods [10]. Researchers seeking to reuse public data must therefore understand also where data are stored, but also the quality and completeness of what they will find.
A systematic evaluation of publicly available scRNA-seq studies from the Gene Expression Omnibus found that only around 40 percent of studies provided readily usable processed count data that could be reliably mapped to repository metadata, and fewer than 10 percent included author-provided cell-type labels [14]. This finding has direct implications for researchers planning to deposit data: the availability of processed data and annotations determines whether others can actually reuse your work.
Core Principles for Database Selection
Data Types and Modalities
Single-cell databases vary in the types of data they accept and host. The most common data type is scRNA-seq, but repositories increasingly accommodate single-cell ATAC sequencing (scATAC-seq), spatial transcriptomics, and multi-omics datasets that measure multiple modalities from the same cells [5]. When selecting a database, confirm that it supports your specific data modality and file formats.
Access Policies and Data Sharing Requirements
Funding agencies increasingly mandate data sharing. The National Institutes of Health Genomic Data Sharing Policy establishes expectations for the deposition and sharing of genomic data generated through NIH-funded research [3]. Researchers should review their funding agreements and institutional requirements before selecting a repository, as some databases are designed to comply with specific policy frameworks.
FAIR Principles
The FAIR Guiding Principles describe the characteristics that make data findable, accessible, interoperable, and reusable [4]. These principles provide a useful framework for evaluating databases. A repository that assigns persistent identifiers, provides clear metadata, uses standard file formats, and documents data provenance will support long-term data reuse more effectively than one that lacks these features.
Integration Capabilities
Because single-cell datasets often need to be combined for meaningful analysis, integration capability is a practical consideration. Benchmarking studies have shown that integration method performance varies substantially depending on the complexity of the integration task and the data modality involved [10]. Some databases provide built-in integration tools or pre-integrated atlases, while others serve primarily as raw data archives.
At a Glance: Major Single-Cell Sequencing Databases
| Database | Primary Data Types | Access Model | Notable Features | Best Suited For |
|---|---|---|---|---|
| Gene Expression Omnibus (GEO) | scRNA-seq, bulk RNA-seq, microarray, spatial transcriptomics | Open access with some controlled access datasets | Extensive metadata, raw and processed data, series and sample organization | General transcriptomics data deposition and retrieval |
| EMBL-EBI Resources | scRNA-seq, multi-omics, proteomics, genomics | Open access | Training materials, cross-referenced databases, European data infrastructure | Researchers seeking integrated European data resources and training |
| NCBI Data Resources | Genomic, transcriptomic, epigenomic data | Open access with controlled access for some datasets | Sequence Read Archive, GEO, dbGaP, comprehensive documentation | Researchers needing integrated genomic data infrastructure |
| Disease-Specific Databases | scRNA-seq, spatial transcriptomics, multi-omics | Open access | Curated disease focus, re-analyzed datasets, interactive visualization | Researchers working on specific diseases or organ systems |
| Atlas Projects | scRNA-seq, scATAC-seq, spatial data | Open access | Large-scale integrated datasets, harmonized cell annotations | Comparative and cross-study analyses |
The table above represents a starting point for database selection. The following sections provide more detailed guidance on specific resources and the practical considerations for using them.
Major General-Purpose Repositories
Gene Expression Omnibus
The Gene Expression Omnibus (GEO) at NCBI functions as a public functional genomics data repository that accepts array-based and sequence-based data [2]. GEO supports MIAME-compliant data submissions and provides tools for querying and downloading data. For single-cell studies, GEO hosts both raw sequencing files and processed count matrices, although the completeness of processed data varies substantially between submissions [14].
When depositing data to GEO, researchers should prepare both raw files and processed count matrices, and should provide cell-type annotations when possible. The finding that fewer than 10 percent of scRNA-seq studies include author-provided cell-type labels suggests that this is a common gap in data deposition practice [14]. Including these annotations substantially increases the reusability of deposited data.
EMBL-EBI Resources
The European Bioinformatics Institute (EMBL-EBI) provides a portfolio of bioinformatics resources that include databases relevant to single-cell research [1]. These resources include the ArrayExpress repository for functional genomics data, the European Nucleotide Archive for raw sequencing data, and the Single Cell Expression Atlas for curated and re-analyzed single-cell expression data.
EMBL-EBI also offers extensive training materials that cover data submission, retrieval, and analysis [1]. For researchers new to single-cell data management, these training resources provide practical guidance on repository selection and data formatting.
NCBI Data Resources
The National Center for Biotechnology Information (NCBI) maintains a comprehensive set of data resources that support single-cell research [2]. Beyond GEO, NCBI provides the Sequence Read Archive (SRA) for raw sequencing data, dbGaP for controlled-access data, and various analysis tools. The integration of these resources allows researchers to trace data from raw sequencing files through to processed expression matrices.
Disease-Specific and Specialized Databases
Respiratory System Database
The scMoresDB database provides a single-cell multi-omics platform specifically tailored for human respiratory diseases [20]. This database re-analyzes single-cell multi-omics datasets and provides cross-omics search capabilities, interactive visualizations, and analytical tools. Example applications have highlighted the potential significance of the BSG receptor in SARS-CoV-2 infection and the involvement of HHIP and TGFB2 in chronic obstructive pulmonary disease [20].
For researchers working on respiratory diseases, this database offers the advantage of curated, re-analyzed data that has been processed with consistent pipelines. This consistency can reduce the burden of data preprocessing and enable more direct cross-study comparisons.
Cancer-Focused Resources
Single-cell sequencing has found extensive application in precision oncology and cancer therapeutics [24]. Cancer-focused databases often integrate single-cell data with clinical information, mutation data, and treatment outcomes. These resources support the identification of cell-type-specific markers and the characterization of tumor microenvironments.
Emerging Specialized Databases
The field continues to produce new specialized databases that address specific data types or research questions. For example, recent work has demonstrated the utility of integrating single-cell RNA sequencing with spatial transcriptomics to study cellular heterogeneity in conditions such as hepatocellular carcinoma [7]. Researchers should monitor the literature for databases relevant to their specific disease or tissue of interest.
Data Integration and Cross-Database Analysis
The Integration Challenge
Single-cell atlases often include samples that span locations, laboratories, and conditions, leading to complex nested batch effects [10]. Joint analysis of atlas datasets requires reliable data integration methods. A comprehensive benchmark of 68 method and preprocessing combinations on 85 batches of gene expression, chromatin accessibility, and simulation data found that highly variable gene selection improves integration performance, while scaling pushes methods to prioritize batch removal over conservation of biological variation [10].
Method Selection
The same benchmark identified scANVI, Scanorama, scVI, and scGen as performing well, particularly on complex integration tasks [10]. However, integration performance for single-cell ATAC sequencing data is strongly affected by the choice of feature space [10]. Researchers should test multiple integration methods on their specific data instead of assuming that a single method will work universally.
Federated Approaches
Privacy concerns can limit data sharing for certain applications. Federated Harmony combines federated learning with the Harmony algorithm to integrate decentralized omics data without raw data sharing [19]. This approach preserves privacy while maintaining integration performance comparable to centralized methods [19]. For researchers working with sensitive data, federated approaches may offer a pathway to collaborative analysis that respects data governance requirements.
Practical Workflow for Data Deposition
Step 1: Review Funding and Institutional Requirements
Before depositing data, review your funding agreements and institutional policies. The NIH Genomic Data Sharing Policy establishes expectations for data sharing that may apply to your research [3]. Determine whether your data requires controlled access or can be made publicly available.
Step 2: Select the Appropriate Repository
Choose a repository that supports your data type and meets any regulatory requirements. General-purpose repositories like GEO and EMBL-EBI resources accept most data types [1][2]. Disease-specific databases may offer additional value through curated re-analysis but may have more restrictive submission criteria.
Step 3: Prepare Complete Metadata
Metadata quality determines data discoverability. Include information about the experimental design, sample preparation, sequencing platform, and analysis pipeline. The FAIR principles emphasize the importance of rich metadata for data findability and reusability [4].
Step 4: Deposit Raw and Processed Data
Deposit both raw sequencing files and processed count matrices. The finding that only around 40 percent of studies provide readily usable processed count data highlights the importance of this step [14]. Processed data should include cell-level expression matrices in standard formats.
Step 5: Provide Cell-Type Annotations
Cell-type annotations substantially increase data reusability, yet fewer than 10 percent of studies include them [14]. If you have performed cell-type annotation as part of your analysis, include these labels in your deposition.
Step 6: Document Analysis Parameters
Document the parameters used for alignment, quantification, and filtering. This documentation enables others to understand how the processed data were generated and to assess whether the data are suitable for their intended analyses.
Practical Workflow for Data Retrieval
Step 1: Define Your Analysis Question
Before searching for data, define the specific question you need to answer. Are you looking for a reference atlas for cell-type annotation? Do you need data from a specific tissue or disease? Are you planning a meta-analysis that requires comparable datasets across studies?
Step 2: Search Multiple Repositories
Search general-purpose repositories like GEO and EMBL-EBI resources, as well as disease-specific databases relevant to your research area [1][2]. Different repositories may host complementary datasets.
Step 3: Assess Data Completeness
Before downloading a dataset, check whether processed count matrices and cell-type annotations are available. The finding that many studies lack readily usable processed data should inform your expectations [14]. If only raw data are available, plan for the computational resources needed to process it.
Step 4: Evaluate Metadata Quality
Review the metadata associated with each dataset. Incomplete or ambiguous metadata can make data interpretation difficult or impossible. The FAIR principles provide a framework for evaluating metadata quality [4].
Step 5: Check for Batch Effects
When combining data from multiple studies, expect batch effects. The benchmark of integration methods found that batch effects are common in atlas-level data and require careful handling [10]. Plan to apply appropriate integration methods and to validate that biological variation is preserved.
Step 6: Document Data Sources
Maintain records of which datasets you used, their accession numbers, and the versions of any processed data. This documentation supports reproducibility and is often required for publication.
Records and Measurements for Data Management
Data Provenance Records
Maintain a data provenance log that records the source of each dataset, the date of download, the accession number, and any processing steps applied. This log supports reproducibility and helps identify data quality issues.
Quality Metrics
Record quality metrics for each dataset you use, including the number of cells, the number of genes detected per cell, the sequencing depth, and any filtering criteria applied. These metrics help you assess whether a dataset is suitable for your analysis.
Integration Validation Records
When integrating data from multiple sources, document the integration method used, the parameters applied, and the validation metrics assessed. The benchmark of integration methods provides guidance on evaluation metrics [10].
Version Control
Track the versions of any software tools used for data processing and analysis. Software updates can change results, so version documentation is essential for reproducibility.
Common Failure Patterns in Database Use
Incomplete Processed Data
The most common barrier to data reuse is the absence of readily usable processed count matrices [14]. Researchers who download data and find only raw sequencing files must invest substantial time in processing before analysis can begin.
Missing Cell-Type Annotations
The scarcity of author-provided cell-type labels limits the immediate utility of many datasets [14]. Researchers may need to perform their own cell-type annotation, which requires reference data and computational resources.
Inadequate Metadata
Poor metadata makes it difficult to determine what a dataset contains and how it was generated. This problem is compounded when researchers attempt to combine datasets from different studies.
Unrecognized Batch Effects
Researchers who combine data from multiple studies without addressing batch effects risk drawing incorrect conclusions. The benchmark of integration methods found that batch effects are complex and require careful handling [10].
Format Incompatibility
Single-cell data are stored in multiple formats, and tools in R and Python may use different data structures. The scDIOR software was developed to address the data transformation problem between platforms, supporting conversion between Seurat, SingleCellExperiment, Monocle, and Scanpy formats [21]. Researchers should plan for format conversion needs.
Limitations and Interpretation Boundaries
mRNA as a Proxy for Protein
Single-cell RNA sequencing measures mRNA abundance, which is often assumed to reflect protein expression. However, post-transcriptional and translational regulation make mRNA an inadequate proxy for protein in many cases [17]. Machine learning methods for protein imputation from scRNA-seq data can improve predictions compared to using cognate mRNAs alone, but these methods require appropriately trained models and highly similar training data [17].
Technical Heterogeneity
Differences in experimental design, sequencing platforms, and sample composition introduce substantial heterogeneity that limits direct comparability between studies [13]. Transcriptomic meta-analysis provides a framework for addressing these challenges, but heterogeneity must be explicitly considered to avoid misleading conclusions [13].
Integration Method Dependence
The choice of integration method can substantially affect results. The benchmark of integration methods found that method performance varies by data type and integration task complexity [10]. Researchers should validate that their conclusions are robust to the choice of integration method.
Generalizability Limits
Machine learning models trained on one dataset may not generalize to other datasets or tissues. The accuracy of protein imputation models depends on the overlap in cell type composition between training and test data [17]. Similar limitations apply to other predictive methods.
Quality Controls and Validation Approaches
Data Quality Assessment
Before analysis, assess the quality of downloaded data. Check the number of cells, genes detected per cell, and sequencing depth. Compare these metrics to published quality standards for the relevant data type.
Integration Validation
When integrating data from multiple sources, validate that the integration preserved biological variation while removing batch effects. The benchmark of integration methods provides a framework for evaluating integration performance [10].
Cross-Study Replication
Validate key findings by checking whether they replicate across independent datasets. Transcriptomic meta-analysis focuses on consistent signals across diverse datasets to enable more robust biological inference [13].
Cell-Type Annotation Validation
If you perform your own cell-type annotation, validate the results using multiple methods. The Augur method provides a machine-learning framework for prioritizing cell types most responsive to biological perturbations, which can complement differential expression analysis [11].
Regulatory and Ethical Considerations
Genomic Data Sharing Policies
The NIH Genomic Data Sharing Policy establishes expectations for the deposition and sharing of genomic data [3]. Researchers should review this policy and any applicable institutional requirements before depositing data.
Privacy and Data Protection
Some single-cell datasets may contain sensitive information, particularly if they include human subjects data. Federated approaches to data integration can help address privacy concerns by avoiding raw data sharing [19].
Data Use Agreements
Some datasets are available under data use agreements that restrict how the data can be used. Review these agreements before downloading and using data.
Publication Requirements
Many journals require data deposition as a condition of publication. Check the requirements of your target journal and select a repository that satisfies those requirements.
Professional Escalation Criteria
When to Seek Expert Assistance
Consider consulting a bioinformatics specialist or data curator when you encounter any of the following situations:
- You need to deposit data that may be subject to controlled access requirements
- You are combining data from many studies with complex batch structures
- You are working with data modalities for which integration methods are less mature
- You need to comply with specific funding agency data sharing requirements
- You are uncertain about the appropriateness of a particular integration method for your data
When to Contact Repository Support
Contact repository support staff when you encounter technical issues with data submission or retrieval. Most repositories provide documentation and support channels for common problems.
When to Escalate Data Quality Concerns
If you identify data quality problems that may affect the interpretation of published results, consider contacting the data depositors or the repository curators. The finding that many datasets lack complete processed data suggests that data quality issues are common [14].
Frequently Asked Questions
What is the difference between raw and processed single-cell data?
Raw data consist of sequencing files in formats such as FASTQ or BAM, which require alignment and quantification before they can be analyzed. Processed data include count matrices that quantify gene expression per cell, often accompanied by cell-level metadata. The availability of processed data varies substantially between studies, with only around 40 percent of scRNA-seq studies providing readily usable processed count matrices [14].
Which database should I use to deposit my single-cell sequencing data?
The choice of database depends on your data type, funding requirements, and research community norms. General-purpose repositories like GEO and EMBL-EBI resources accept most data types and are widely recognized [1][2]. Disease-specific databases may offer additional value through curated re-analysis but may have more restrictive submission criteria [20]. Review your funding agreements and journal requirements before selecting a repository.
How do I find single-cell data for a specific tissue or disease?
Search general-purpose repositories like GEO and EMBL-EBI resources using keywords related to your tissue or disease of interest [1][2]. Also search disease-specific databases, such as scMoresDB for respiratory diseases [20]. Review the metadata of candidate datasets to assess data completeness and quality.
What should I do if a dataset only has raw sequencing files?
If only raw data are available, you will need to process the data yourself. This requires alignment, quantification, and quality filtering. Plan for the computational resources and time needed for this processing. The finding that many studies lack readily usable processed data should inform your expectations [14].
How do I combine data from multiple single-cell studies?
Combining data from multiple studies requires careful attention to batch effects. Benchmarking studies have shown that integration method performance varies by data type and task complexity [10]. Test multiple integration methods and validate that biological variation is preserved. Document your integration approach and validation metrics.
Are cell-type annotations available for most public datasets?
No. Fewer than 10 percent of scRNA-seq studies include author-provided cell-type labels [14]. If annotations are not available, you may need to perform your own cell-type annotation using reference data and computational tools.
What are the data sharing requirements for NIH-funded research?
The NIH Genomic Data Sharing Policy establishes expectations for the deposition and sharing of genomic data generated through NIH-funded research [3]. Review this policy and your specific funding agreement to understand your obligations.
How can I ensure my deposited data will be reusable by others?
Deposit both raw and processed data, provide complete metadata, and include cell-type annotations when possible. The finding that many datasets lack processed data and annotations highlights the importance of these practices [14]. Following the FAIR principles can help ensure your data are findable, accessible, interoperable, and reusable [4].
Related Bioinformatics Guides
- Single-Cell RNA Sequencing: From Bulk to Resolution
- Master Guide: Single-Cell RNA Sequencing Bioinformatics Workflows
- Foundation Models for Single-Cell Biology
- Single-Cell ATAC-Seq Bioinformatics
- Data Sharing and Privacy in Genomic Research
References and Further Reading
- EMBL-EBI Training. European Bioinformatics Institute.
- NCBI Data Resources. National Center for Biotechnology Information.
- Genomic Data Sharing Policy. National Institutes of Health.
- The FAIR Guiding Principles. Scientific Data.
- Comprehensive Integration of Single-Cell Data.. Cell, 2019.
- mRNA-Seq whole-transcriptome analysis of a single cell.. Nature methods, 2009.
- Hepatic stellate cells promote hepatocellular carcinoma development by regulating histone lactylation: Novel insights from single-cell RNA sequencing and spatial transcriptomics analyses.. Cancer letters, 2024.
- Single-Cell RNA Sequencing Unveils Unique Transcriptomic Signatures of Organ-Specific Endothelial Cells.. Circulation, 2020.
- Integrative single-cell RNA sequencing and mendelian randomization analysis reveal the potential role of synaptic vesicle cycling-related genes in Alzheimer's disease.. The journal of prevention of Alzheimer's disease, 2025.
- Benchmarking atlas-level data integration in single-cell genomics.. Nature methods, 2022.
- Cell type prioritization in single-cell data.. Nature biotechnology, 2021.
- Causal effects of gut microbiota on sepsis and sepsis-related death: insights from genome-wide Mendelian randomization, single-cell RNA, bulk RNA sequencing, and network pharmacology.. Journal of translational medicine, 2024.
- Transcriptomic Meta-Analysis as a Framework for Robust Cross-Study Biological Inference.. 2026.
- Persistent hindrances to data re-use in single-cell genomics.. 2026.
- Artificial Intelligence in Transcriptomics: From Human-in-the-Loop to Agentic AI.. 2026.
- Tracing the stemness and malignant transition in a heritable colorectal cancer Lynch Syndrome by single-cell RNA-seq analysis
- Machine learning predictions surpass individual mRNAs as a proxy of single-cell protein expression.. 2026.
- HistoMap: Reconstructing Spatially Resolved Single-Cell Profiles from Bulk RNA-Seq to Decipher the Immune-Excluded Microenvironment in Colon Cancer. 2026.
- Harmony-based data integration for distributed single-cell multi-omics data. PLoS Comput. Biol., 2025.
- scMoresDB: A comprehensive database of single-cell multi-omics data for human respiratory system. iScience, 2024.
- scDIOR: single cell RNA-seq data IO software. BMC Bioinformatics, 2022.
- Integrated single-cell RNA sequencing and multi-database analysis reveal THBS2 promotes high glucose-induced mesangial cell oxidative stress and fibrosis in diabetic nephropathy. Electronic Journal of Biotechnology, 2026.
- Navigating single-cell RNA-sequencing: protocols, tools, databases, and applications. Genomics and Informatics, 2025.
- Single-Cell Sequencing: Current Applications in Precision Onco-Genomics and Cancer Therapeutics. Cancers, 2022.
- Application of single-cell RNA sequencing technology in the study of skeletal system biology. Chinese Journal of Tissue Engineering Research, 2021.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.