Zubair Khalid

Virologist/Molecular Biologist | Veterinarian | Bioinformatician

Conventional & Molecular Virology • Vaccine Development • Computational Biology

Dr. Zubair Khalid is a veterinarian and virologist specializing in conventional and molecular virology, vaccine development, and computational biology. Dedicated to advancing animal health through innovative research and multi-omics approaches.

Dr. Zubair Khalid - Veterinarian, Virologist, and Vaccine Development Researcher specializing in Computational Biology, Multi-omics, Animal Health, and Infectious Disease Research

Section: Infrastructure, Cloud & Policy

Single-Cell Sequencing Databases: Resources for Data Sharing and Exploration

Single-cell sequencing has transformed biological research by enabling transcriptome analysis at the resolution of individual cells. The first single-cell mRNA sequencing assays demonstrated that a single mouse blastomere could yield detection of thousands more genes than microarray techniques, revealing transcript isoform complexity that bulk methods could not resolve [6]. As these technologies have matured, the volume of publicly available single-cell data has grown exponentially, creating both opportunity and challenge for researchers who need to locate, access, and reuse datasets. This article provides a practical orientation to the major databases and repositories for single-cell sequencing data, with attention to access policies, data formats, and the practical decisions researchers face when depositing or retrieving data.

The Data Landscape in Single-Cell Genomics

Single-cell RNA sequencing (scRNA-seq) has become a standard tool for characterizing cell states and heterogeneity, but the analytical value of any single experiment depends on the ability to compare findings against existing data [5]. Public repositories serve as the primary infrastructure for this comparison, hosting raw sequencing files, processed count matrices, and increasingly, cell-type annotations and spatial transcriptomics data.

The scale of available data presents both opportunities and complications. Large atlas projects aggregate samples across laboratories, locations, and experimental conditions, introducing complex batch effects that require careful integration methods [10]. Researchers seeking to reuse public data must therefore understand also where data are stored, but also the quality and completeness of what they will find.

A systematic evaluation of publicly available scRNA-seq studies from the Gene Expression Omnibus found that only around 40 percent of studies provided readily usable processed count data that could be reliably mapped to repository metadata, and fewer than 10 percent included author-provided cell-type labels [14]. This finding has direct implications for researchers planning to deposit data: the availability of processed data and annotations determines whether others can actually reuse your work.

Core Principles for Database Selection

Data Types and Modalities

Single-cell databases vary in the types of data they accept and host. The most common data type is scRNA-seq, but repositories increasingly accommodate single-cell ATAC sequencing (scATAC-seq), spatial transcriptomics, and multi-omics datasets that measure multiple modalities from the same cells [5]. When selecting a database, confirm that it supports your specific data modality and file formats.

Access Policies and Data Sharing Requirements

Funding agencies increasingly mandate data sharing. The National Institutes of Health Genomic Data Sharing Policy establishes expectations for the deposition and sharing of genomic data generated through NIH-funded research [3]. Researchers should review their funding agreements and institutional requirements before selecting a repository, as some databases are designed to comply with specific policy frameworks.

FAIR Principles

The FAIR Guiding Principles describe the characteristics that make data findable, accessible, interoperable, and reusable [4]. These principles provide a useful framework for evaluating databases. A repository that assigns persistent identifiers, provides clear metadata, uses standard file formats, and documents data provenance will support long-term data reuse more effectively than one that lacks these features.

Integration Capabilities

Because single-cell datasets often need to be combined for meaningful analysis, integration capability is a practical consideration. Benchmarking studies have shown that integration method performance varies substantially depending on the complexity of the integration task and the data modality involved [10]. Some databases provide built-in integration tools or pre-integrated atlases, while others serve primarily as raw data archives.

At a Glance: Major Single-Cell Sequencing Databases

Database Primary Data Types Access Model Notable Features Best Suited For
Gene Expression Omnibus (GEO) scRNA-seq, bulk RNA-seq, microarray, spatial transcriptomics Open access with some controlled access datasets Extensive metadata, raw and processed data, series and sample organization General transcriptomics data deposition and retrieval
EMBL-EBI Resources scRNA-seq, multi-omics, proteomics, genomics Open access Training materials, cross-referenced databases, European data infrastructure Researchers seeking integrated European data resources and training
NCBI Data Resources Genomic, transcriptomic, epigenomic data Open access with controlled access for some datasets Sequence Read Archive, GEO, dbGaP, comprehensive documentation Researchers needing integrated genomic data infrastructure
Disease-Specific Databases scRNA-seq, spatial transcriptomics, multi-omics Open access Curated disease focus, re-analyzed datasets, interactive visualization Researchers working on specific diseases or organ systems
Atlas Projects scRNA-seq, scATAC-seq, spatial data Open access Large-scale integrated datasets, harmonized cell annotations Comparative and cross-study analyses

The table above represents a starting point for database selection. The following sections provide more detailed guidance on specific resources and the practical considerations for using them.

Major General-Purpose Repositories

Gene Expression Omnibus

The Gene Expression Omnibus (GEO) at NCBI functions as a public functional genomics data repository that accepts array-based and sequence-based data [2]. GEO supports MIAME-compliant data submissions and provides tools for querying and downloading data. For single-cell studies, GEO hosts both raw sequencing files and processed count matrices, although the completeness of processed data varies substantially between submissions [14].

When depositing data to GEO, researchers should prepare both raw files and processed count matrices, and should provide cell-type annotations when possible. The finding that fewer than 10 percent of scRNA-seq studies include author-provided cell-type labels suggests that this is a common gap in data deposition practice [14]. Including these annotations substantially increases the reusability of deposited data.

EMBL-EBI Resources

The European Bioinformatics Institute (EMBL-EBI) provides a portfolio of bioinformatics resources that include databases relevant to single-cell research [1]. These resources include the ArrayExpress repository for functional genomics data, the European Nucleotide Archive for raw sequencing data, and the Single Cell Expression Atlas for curated and re-analyzed single-cell expression data.

EMBL-EBI also offers extensive training materials that cover data submission, retrieval, and analysis [1]. For researchers new to single-cell data management, these training resources provide practical guidance on repository selection and data formatting.

NCBI Data Resources

The National Center for Biotechnology Information (NCBI) maintains a comprehensive set of data resources that support single-cell research [2]. Beyond GEO, NCBI provides the Sequence Read Archive (SRA) for raw sequencing data, dbGaP for controlled-access data, and various analysis tools. The integration of these resources allows researchers to trace data from raw sequencing files through to processed expression matrices.

Disease-Specific and Specialized Databases

Respiratory System Database

The scMoresDB database provides a single-cell multi-omics platform specifically tailored for human respiratory diseases [20]. This database re-analyzes single-cell multi-omics datasets and provides cross-omics search capabilities, interactive visualizations, and analytical tools. Example applications have highlighted the potential significance of the BSG receptor in SARS-CoV-2 infection and the involvement of HHIP and TGFB2 in chronic obstructive pulmonary disease [20].

For researchers working on respiratory diseases, this database offers the advantage of curated, re-analyzed data that has been processed with consistent pipelines. This consistency can reduce the burden of data preprocessing and enable more direct cross-study comparisons.

Cancer-Focused Resources

Single-cell sequencing has found extensive application in precision oncology and cancer therapeutics [24]. Cancer-focused databases often integrate single-cell data with clinical information, mutation data, and treatment outcomes. These resources support the identification of cell-type-specific markers and the characterization of tumor microenvironments.

Emerging Specialized Databases

The field continues to produce new specialized databases that address specific data types or research questions. For example, recent work has demonstrated the utility of integrating single-cell RNA sequencing with spatial transcriptomics to study cellular heterogeneity in conditions such as hepatocellular carcinoma [7]. Researchers should monitor the literature for databases relevant to their specific disease or tissue of interest.

Data Integration and Cross-Database Analysis

The Integration Challenge

Single-cell atlases often include samples that span locations, laboratories, and conditions, leading to complex nested batch effects [10]. Joint analysis of atlas datasets requires reliable data integration methods. A comprehensive benchmark of 68 method and preprocessing combinations on 85 batches of gene expression, chromatin accessibility, and simulation data found that highly variable gene selection improves integration performance, while scaling pushes methods to prioritize batch removal over conservation of biological variation [10].

Method Selection

The same benchmark identified scANVI, Scanorama, scVI, and scGen as performing well, particularly on complex integration tasks [10]. However, integration performance for single-cell ATAC sequencing data is strongly affected by the choice of feature space [10]. Researchers should test multiple integration methods on their specific data instead of assuming that a single method will work universally.

Federated Approaches

Privacy concerns can limit data sharing for certain applications. Federated Harmony combines federated learning with the Harmony algorithm to integrate decentralized omics data without raw data sharing [19]. This approach preserves privacy while maintaining integration performance comparable to centralized methods [19]. For researchers working with sensitive data, federated approaches may offer a pathway to collaborative analysis that respects data governance requirements.

Practical Workflow for Data Deposition

Step 1: Review Funding and Institutional Requirements

Before depositing data, review your funding agreements and institutional policies. The NIH Genomic Data Sharing Policy establishes expectations for data sharing that may apply to your research [3]. Determine whether your data requires controlled access or can be made publicly available.

Step 2: Select the Appropriate Repository

Choose a repository that supports your data type and meets any regulatory requirements. General-purpose repositories like GEO and EMBL-EBI resources accept most data types [1][2]. Disease-specific databases may offer additional value through curated re-analysis but may have more restrictive submission criteria.

Step 3: Prepare Complete Metadata

Metadata quality determines data discoverability. Include information about the experimental design, sample preparation, sequencing platform, and analysis pipeline. The FAIR principles emphasize the importance of rich metadata for data findability and reusability [4].

Step 4: Deposit Raw and Processed Data

Deposit both raw sequencing files and processed count matrices. The finding that only around 40 percent of studies provide readily usable processed count data highlights the importance of this step [14]. Processed data should include cell-level expression matrices in standard formats.

Step 5: Provide Cell-Type Annotations

Cell-type annotations substantially increase data reusability, yet fewer than 10 percent of studies include them [14]. If you have performed cell-type annotation as part of your analysis, include these labels in your deposition.

Step 6: Document Analysis Parameters

Document the parameters used for alignment, quantification, and filtering. This documentation enables others to understand how the processed data were generated and to assess whether the data are suitable for their intended analyses.

Practical Workflow for Data Retrieval

Step 1: Define Your Analysis Question

Before searching for data, define the specific question you need to answer. Are you looking for a reference atlas for cell-type annotation? Do you need data from a specific tissue or disease? Are you planning a meta-analysis that requires comparable datasets across studies?

Step 2: Search Multiple Repositories

Search general-purpose repositories like GEO and EMBL-EBI resources, as well as disease-specific databases relevant to your research area [1][2]. Different repositories may host complementary datasets.

Step 3: Assess Data Completeness

Before downloading a dataset, check whether processed count matrices and cell-type annotations are available. The finding that many studies lack readily usable processed data should inform your expectations [14]. If only raw data are available, plan for the computational resources needed to process it.

Step 4: Evaluate Metadata Quality

Review the metadata associated with each dataset. Incomplete or ambiguous metadata can make data interpretation difficult or impossible. The FAIR principles provide a framework for evaluating metadata quality [4].

Step 5: Check for Batch Effects

When combining data from multiple studies, expect batch effects. The benchmark of integration methods found that batch effects are common in atlas-level data and require careful handling [10]. Plan to apply appropriate integration methods and to validate that biological variation is preserved.

Step 6: Document Data Sources

Maintain records of which datasets you used, their accession numbers, and the versions of any processed data. This documentation supports reproducibility and is often required for publication.

Records and Measurements for Data Management

Data Provenance Records

Maintain a data provenance log that records the source of each dataset, the date of download, the accession number, and any processing steps applied. This log supports reproducibility and helps identify data quality issues.

Quality Metrics

Record quality metrics for each dataset you use, including the number of cells, the number of genes detected per cell, the sequencing depth, and any filtering criteria applied. These metrics help you assess whether a dataset is suitable for your analysis.

Integration Validation Records

When integrating data from multiple sources, document the integration method used, the parameters applied, and the validation metrics assessed. The benchmark of integration methods provides guidance on evaluation metrics [10].

Version Control

Track the versions of any software tools used for data processing and analysis. Software updates can change results, so version documentation is essential for reproducibility.

Common Failure Patterns in Database Use

Incomplete Processed Data

The most common barrier to data reuse is the absence of readily usable processed count matrices [14]. Researchers who download data and find only raw sequencing files must invest substantial time in processing before analysis can begin.

Missing Cell-Type Annotations

The scarcity of author-provided cell-type labels limits the immediate utility of many datasets [14]. Researchers may need to perform their own cell-type annotation, which requires reference data and computational resources.

Inadequate Metadata

Poor metadata makes it difficult to determine what a dataset contains and how it was generated. This problem is compounded when researchers attempt to combine datasets from different studies.

Unrecognized Batch Effects

Researchers who combine data from multiple studies without addressing batch effects risk drawing incorrect conclusions. The benchmark of integration methods found that batch effects are complex and require careful handling [10].

Format Incompatibility

Single-cell data are stored in multiple formats, and tools in R and Python may use different data structures. The scDIOR software was developed to address the data transformation problem between platforms, supporting conversion between Seurat, SingleCellExperiment, Monocle, and Scanpy formats [21]. Researchers should plan for format conversion needs.

Limitations and Interpretation Boundaries

mRNA as a Proxy for Protein

Single-cell RNA sequencing measures mRNA abundance, which is often assumed to reflect protein expression. However, post-transcriptional and translational regulation make mRNA an inadequate proxy for protein in many cases [17]. Machine learning methods for protein imputation from scRNA-seq data can improve predictions compared to using cognate mRNAs alone, but these methods require appropriately trained models and highly similar training data [17].

Technical Heterogeneity

Differences in experimental design, sequencing platforms, and sample composition introduce substantial heterogeneity that limits direct comparability between studies [13]. Transcriptomic meta-analysis provides a framework for addressing these challenges, but heterogeneity must be explicitly considered to avoid misleading conclusions [13].

Integration Method Dependence

The choice of integration method can substantially affect results. The benchmark of integration methods found that method performance varies by data type and integration task complexity [10]. Researchers should validate that their conclusions are robust to the choice of integration method.

Generalizability Limits

Machine learning models trained on one dataset may not generalize to other datasets or tissues. The accuracy of protein imputation models depends on the overlap in cell type composition between training and test data [17]. Similar limitations apply to other predictive methods.

Quality Controls and Validation Approaches

Data Quality Assessment

Before analysis, assess the quality of downloaded data. Check the number of cells, genes detected per cell, and sequencing depth. Compare these metrics to published quality standards for the relevant data type.

Integration Validation

When integrating data from multiple sources, validate that the integration preserved biological variation while removing batch effects. The benchmark of integration methods provides a framework for evaluating integration performance [10].

Cross-Study Replication

Validate key findings by checking whether they replicate across independent datasets. Transcriptomic meta-analysis focuses on consistent signals across diverse datasets to enable more robust biological inference [13].

Cell-Type Annotation Validation

If you perform your own cell-type annotation, validate the results using multiple methods. The Augur method provides a machine-learning framework for prioritizing cell types most responsive to biological perturbations, which can complement differential expression analysis [11].

Regulatory and Ethical Considerations

Genomic Data Sharing Policies

The NIH Genomic Data Sharing Policy establishes expectations for the deposition and sharing of genomic data [3]. Researchers should review this policy and any applicable institutional requirements before depositing data.

Privacy and Data Protection

Some single-cell datasets may contain sensitive information, particularly if they include human subjects data. Federated approaches to data integration can help address privacy concerns by avoiding raw data sharing [19].

Data Use Agreements

Some datasets are available under data use agreements that restrict how the data can be used. Review these agreements before downloading and using data.

Publication Requirements

Many journals require data deposition as a condition of publication. Check the requirements of your target journal and select a repository that satisfies those requirements.

Professional Escalation Criteria

When to Seek Expert Assistance

Consider consulting a bioinformatics specialist or data curator when you encounter any of the following situations:

  • You need to deposit data that may be subject to controlled access requirements
  • You are combining data from many studies with complex batch structures
  • You are working with data modalities for which integration methods are less mature
  • You need to comply with specific funding agency data sharing requirements
  • You are uncertain about the appropriateness of a particular integration method for your data

When to Contact Repository Support

Contact repository support staff when you encounter technical issues with data submission or retrieval. Most repositories provide documentation and support channels for common problems.

When to Escalate Data Quality Concerns

If you identify data quality problems that may affect the interpretation of published results, consider contacting the data depositors or the repository curators. The finding that many datasets lack complete processed data suggests that data quality issues are common [14].

Frequently Asked Questions

What is the difference between raw and processed single-cell data?

Raw data consist of sequencing files in formats such as FASTQ or BAM, which require alignment and quantification before they can be analyzed. Processed data include count matrices that quantify gene expression per cell, often accompanied by cell-level metadata. The availability of processed data varies substantially between studies, with only around 40 percent of scRNA-seq studies providing readily usable processed count matrices [14].

Which database should I use to deposit my single-cell sequencing data?

The choice of database depends on your data type, funding requirements, and research community norms. General-purpose repositories like GEO and EMBL-EBI resources accept most data types and are widely recognized [1][2]. Disease-specific databases may offer additional value through curated re-analysis but may have more restrictive submission criteria [20]. Review your funding agreements and journal requirements before selecting a repository.

How do I find single-cell data for a specific tissue or disease?

Search general-purpose repositories like GEO and EMBL-EBI resources using keywords related to your tissue or disease of interest [1][2]. Also search disease-specific databases, such as scMoresDB for respiratory diseases [20]. Review the metadata of candidate datasets to assess data completeness and quality.

What should I do if a dataset only has raw sequencing files?

If only raw data are available, you will need to process the data yourself. This requires alignment, quantification, and quality filtering. Plan for the computational resources and time needed for this processing. The finding that many studies lack readily usable processed data should inform your expectations [14].

How do I combine data from multiple single-cell studies?

Combining data from multiple studies requires careful attention to batch effects. Benchmarking studies have shown that integration method performance varies by data type and task complexity [10]. Test multiple integration methods and validate that biological variation is preserved. Document your integration approach and validation metrics.

Are cell-type annotations available for most public datasets?

No. Fewer than 10 percent of scRNA-seq studies include author-provided cell-type labels [14]. If annotations are not available, you may need to perform your own cell-type annotation using reference data and computational tools.

What are the data sharing requirements for NIH-funded research?

The NIH Genomic Data Sharing Policy establishes expectations for the deposition and sharing of genomic data generated through NIH-funded research [3]. Review this policy and your specific funding agreement to understand your obligations.

How can I ensure my deposited data will be reusable by others?

Deposit both raw and processed data, provide complete metadata, and include cell-type annotations when possible. The finding that many datasets lack processed data and annotations highlights the importance of these practices [14]. Following the FAIR principles can help ensure your data are findable, accessible, interoperable, and reusable [4].

Related Bioinformatics Guides

References and Further Reading

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.