How to Access Single-Cell RNA-seq Data from Public Repositories: GEO, SRA, and Beyond
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Repository Specialization: GEO primarily hosts processed gene expression matrices and curated metadata, serving as a practical starting point for downstream analysis. SRA stores raw sequencing reads, essential for full pipeline reproduction or custom alignment, but requires significant computational effort for processing into count matrices.
- Metadata is Foundational: The accuracy and completeness of metadata in repositories like GEO are critical for dataset selection and biological interpretation. Key metadata elements include organism, tissue source, experimental conditions, platform, and cell count, which directly impact the validity of downstream analyses.
- File Format Dictates Tool Compatibility: 10x Genomics data, commonly distributed in MEX or H5 formats, is directly readable by popular tools like Seurat and Scanpy. Other platforms may yield per-cell files or CSV/TSV matrices, necessitating format conversion or specialized parsing before analysis.
- Raw vs. Processed Data Tradeoff: Accessing raw FASTQ files from SRA offers maximum analytical control but demands substantial computational resources for alignment and quantification. Processed count matrices from GEO bypass this computational burden but inherit the original authors' processing choices, requiring careful documentation.
- Integrated Tools Streamline Access: Specialized platforms like Celline automate data retrieval from multiple repositories, metadata extraction using LLMs, and integration of established analysis tools (e.g., Scrublet for doublet removal, Seurat/Scanpy for QC, Harmony for batch correction). This significantly reduces manual curation effort for large-scale multi-dataset projects.
- Quality Control is Paramount: Post-download validation is essential, including verifying cell counts against reported metrics, assessing technical quality indicators (e.g., UMIs per cell, mitochondrial read percentage), and checking for artifacts like doublets (using tools like Scrublet or scUmaper) and ambient RNA contamination.
Single-cell RNA sequencing (scRNA-seq) generates transcriptome profiles for individual cells, providing biological resolution that bulk RNA sequencing cannot match, but at the cost of increased technical noise and data complexity. Accessing public scRNA-seq data requires navigating repositories such as the Gene Expression Omnibus (GEO), the Sequence Read Archive (SRA), and specialized single-cell atlases, then converting raw or processed files into analysis-ready formats. This article provides a practical workflow for locating, downloading, and validating scRNA-seq datasets, with attention to file formats, metadata quality, and reproducibility.
Understanding the Public Data Landscape for Single-Cell Transcriptomics
Public repositories for scRNA-seq data differ in what they store, how they organize accessions, and which file formats they support. The National Center for Biotechnology Information (NCBI) hosts GEO for curated functional genomics data and SRA for raw sequencing reads, alongside associated search and analysis services. The European Bioinformatics Institute (EMBL-EBI) provides complementary training pathways and data resources that help researchers build the skills needed to work with these repositories. Understanding the division of labor between these archives is the first step in designing a data access strategy.
GEO serves as the primary entry point for processed and curated datasets, including gene expression matrices, metadata, and supplementary files. Most published scRNA-seq studies deposit their processed count matrices in GEO, often alongside the raw FASTQ files in SRA. The SRA stores raw sequencing reads in compressed formats and requires additional processing to generate count matrices. For researchers who want to avoid the computational burden of aligning raw reads, GEO processed data is usually the more practical starting point.
Specialized resources extend beyond these core archives. The Human Cell Atlas and similar projects maintain their own data portals, while tools such as Celline provide one-step retrieval and integrative analysis of public scRNA-seq data from multiple repositories. Celline automatically gathers raw data, extracts metadata using large language models, and wraps established tools including Scrublet for doublet removal, Seurat and Scanpy for quality control and cell-type annotation, Harmony and scVI for batch correction, and Slingshot for trajectory inference. This type of integrated tool reduces the manual curation burden that typically accompanies multi-repository data access.
At a Glance: Repository Selection and File Format Decisions
| Data Need | Recommended Repository | File Format | Processing Required | Best Use Case |
|---|---|---|---|---|
| Processed count matrix | GEO | Matrix Market, CSV, H5 | Minimal, load into Seurat or Scanpy | Downstream analysis, cell typing, differential expression |
| Raw sequencing reads | SRA | FASTQ, SRA format | Alignment, quantification, count matrix generation | Reproducing full pipeline, custom alignment parameters |
| 10x Genomics output | GEO supplementary files or 10x Cloud | H5, MEX (matrix, features, barcodes) | Load directly into Cell Ranger output readers | Standard scRNA-seq analysis with Cell Ranger-derived data |
| Multi-modal or atlas data | EMBL-EBI resources, Human Cell Atlas | H5AD, H5, loom | Format conversion, integration | Cross-dataset integration, weakly linked modality analysis |
| Automated retrieval and integration | Celline or similar tools | Multiple formats | One-line commands per step | Large-scale multi-dataset projects |
Core Principles of Single-Cell Data Access
Metadata Is the Foundation of Dataset Selection
The quality of any downstream analysis depends on the accuracy and completeness of dataset metadata. GEO records include study design, organism, tissue, platform, and sample characteristics, but the level of detail varies substantially between submissions. Published studies frequently report the GEO accession numbers they used, which provides a direct path to the underlying data. For example, studies of acute myocardial infarction used GSE180678 for scRNA-seq data and GSE182923 for gene expression data, while COPD research used GSE173896 for single-cell data and GSE57148 for bulk RNA-seq. These accession numbers appear in the methods sections of published papers and serve as reliable entry points.
When evaluating a dataset, check whether the metadata describes the single-cell platform, the number of cells, the sequencing depth, and the tissue source. Datasets that lack this information may still be usable, but the uncertainty propagates into downstream interpretation. The NCBI search systems allow filtering by organism, study type, and platform, which narrows the search space considerably.
Raw Data versus Processed Data: A Tradeoff in Control and Effort
Raw sequencing reads in SRA provide maximum control over the analysis pipeline. Researchers can choose their own alignment software, quantify against their preferred reference genome, and apply custom quality filters. This control comes at a computational cost. A typical scRNA-seq dataset contains millions of reads per cell, and aligning these reads requires substantial compute resources and bioinformatics expertise.
Processed data in GEO bypasses the alignment step but inherits the processing choices of the original authors. The count matrix reflects their choice of aligner, reference genome version, and quality thresholds. For most downstream analyses, including cell typing, differential expression, and trajectory inference, processed data is sufficient. The key is to document which processing decisions were made by the original authors and to cite the dataset accordingly.
File Formats Determine Tool Compatibility
The 10x Genomics platform produces a specific set of output files that have become a de facto standard in the field. The MEX format includes three files: a matrix of counts, a features file listing genes, and a barcodes file identifying cells. The H5 format packages these components into a single hierarchical file. Seurat and Scanpy both provide functions to read these formats directly, which simplifies the transition from download to analysis.
Other platforms produce different formats. Smart-seq2 data often arrives as per-cell files or as a single count matrix in CSV or TSV format. Drop-seq data may require additional parsing. Before downloading a dataset, check the supplementary files to confirm that the format matches the tools you plan to use. The Bioconductor project provides packages and workflows for low-level analysis of scRNA-seq data, including quality control, normalization, clustering, and marker gene detection, and these packages have specific input format requirements.
Practical Workflow for Locating and Downloading scRNA-seq Data
Step 1: Define the Search Strategy
Start with the biological question and work backward to the data. If the goal is to study a specific disease or tissue, search GEO using disease terms, tissue terms, and the phrase "single-cell RNA-seq" or "scRNA-seq." Published papers in the field provide a faster route, since their methods sections list the exact accession numbers used. For example, studies of intestinal-type gastric cancer downloaded scRNA-seq data from GEO and used TCGA data for validation, while colorectal cancer research combined GEO scRNA-seq data with bulk RNA-seq from UCSC Xena.
The NCBI search interface supports boolean operators and field-specific queries. A search for "single-cell RNA-seq[Title] AND disease[Title]" returns studies with those terms in the title, which often indicates a focused dataset. Filtering by organism and platform further narrows the results.
Step 2: Evaluate Dataset Suitability
Before downloading, examine the GEO record to assess whether the dataset meets the needs of the analysis. Key questions include:
- How many cells are profiled and from which conditions?
- What is the sequencing platform and depth?
- Are the raw data, processed data, or both available?
- What supplementary files are provided and in which formats?
- Is there sufficient metadata to interpret cell types and experimental conditions?
The GEO record includes a summary, overall design, and links to supplementary files. For scRNA-seq datasets, the supplementary files typically include the count matrix and often the Cell Ranger output files. Downloading the processed data first and inspecting the matrix dimensions provides a quick check on whether the dataset matches the description.
Step 3: Download Processed Data from GEO
Processed data downloads from GEO use HTTPS or FTP. The supplementary files section lists each file with its size and format. For 10x Genomics data, download the filtered feature-barcode matrix files, which contain only the cells that passed the Cell Ranger filtering. The raw feature-barcode matrix includes all barcodes and requires additional filtering.
Use a download manager or command-line tool such as wget or curl for large files. Verify file integrity by comparing checksums if the repository provides them. After download, inspect the file structure before loading into analysis software.
Step 4: Download Raw Data from SRA When Needed
Raw data downloads from SRA require the SRA Toolkit, which includes the prefetch and fasterq-dump commands. The SRA Toolkit converts the compressed SRA format to FASTQ files that can be used with alignment software. For 10x Genomics data, the FASTQ files follow a naming convention that includes the sample index and read type, which Cell Ranger uses to organize the data.
Raw data downloads are substantially larger than processed data. A single scRNA-seq sample can require tens of gigabytes of storage. Plan for sufficient disk space and bandwidth before starting the download. The SRA provides access to data through cloud-based endpoints, which can improve download speeds for large datasets.
Step 5: Load Data into Analysis Software
Seurat, an R package, provides Read10X functions for 10x Genomics output and general read functions for other matrix formats. Scanpy, a Python package, provides similar functionality. The Bioconductor project offers additional packages for single-cell analysis, including tools for quality control, normalization, and clustering. The choice of software depends on the analysis goals and the researcher's programming language preference.
When loading data, verify that the gene identifiers match the reference genome used in the analysis. Some datasets use gene symbols, while others use Ensembl IDs or Entrez IDs. Conversion between identifier types may be necessary, and this step should be documented in the analysis record.
Options and Tradeoffs in Data Access Tools
GEOquery and Programmatic Access
The GEOquery package for R provides programmatic access to GEO records, allowing researchers to download series matrices and supplementary files from within an R session. This approach supports reproducible workflows, since the download commands can be included in analysis scripts. GEOquery works well for processed data but does not handle raw sequencing reads, which require the SRA Toolkit.
Cell Ranger and 10x Genomics Tools
Cell Ranger is the standard pipeline for processing 10x Genomics data, converting raw FASTQ files into count matrices. Researchers who download raw data from SRA must run Cell Ranger or an equivalent alignment pipeline before they can use the data in Seurat or Scanpy. The Cell Ranger output includes filtered and unfiltered matrices, which provides flexibility in quality control decisions.
Integrated Retrieval Tools
Tools such as Celline address the fragmentation of public scRNA-seq resources by providing an end-to-end workflow for retrieval, preprocessing, integration, and analysis. Celline automatically gathers raw data from multiple repositories, extracts metadata, and wraps established analysis tools into one-line commands. This approach reduces the manual effort of curating accessions and metadata, particularly for projects that use multiple datasets. The tradeoff is that integrated tools may not offer the same level of control as manual workflows, and researchers should verify that the tool's processing choices match their analysis needs.
Cloud-Based Access
Some repositories provide cloud-based access to data, allowing researchers to run analyses without downloading large files to local storage. The NCBI and EMBL-EBI both support cloud-based data access through partnerships with cloud providers. This approach works well for researchers who have access to cloud computing resources and prefer to avoid local storage requirements.
Quality Control and Data Validation After Download
Verify Cell Counts and Gene Detection
After loading a dataset, check that the number of cells and genes matches the GEO record description. Discrepancies may indicate incomplete downloads, incorrect file selection, or differences in filtering between the original analysis and the deposited data. Published studies report these metrics, providing a reference for validation. For example, one endometrial carcinoma study detected 33,408 genes in 33,162 cells from scRNA-seq data, and these numbers can be compared against the downloaded matrix.
Assess Technical Quality Metrics
Standard scRNA-seq quality metrics include the number of unique molecular identifiers (UMIs) per cell, the number of genes detected per cell, and the percentage of mitochondrial reads. These metrics identify low-quality cells and doublets. The Seurat package provides functions to calculate and visualize these metrics, and the Bioconductor workflow for low-level analysis covers quality control in detail. The sctransform package offers a modeling framework for normalization and variance stabilization that removes the influence of technical characteristics while preserving biological heterogeneity.
Check for Doublets and Ambient RNA
Doublets, where two cells are captured in the same droplet, are a common artifact in scRNA-seq data. Tools such as Scrublet and scUmaper identify doublets using simulation-based and biologically grounded approaches. scUmaper codifies lineage-marker incompatibility rules and applies global clustering followed by within-lineage re-clustering to reveal anomalous subclusters with implausible cross-lineage co-expression. Across six public human organ datasets, scUmaper removed additional high-confidence heterotypic doublets that were retained by simulation-based approaches.
Ambient RNA contamination, where RNA from lysed cells is captured in droplets, can also distort expression profiles. The Cell Ranger pipeline estimates ambient RNA profiles and provides a CellBender or SoupX correction step. Researchers should document whether ambient RNA correction was applied and which method was used.
Validate Cell Type Annotations
Cell type annotations from the original study provide a reference for validation. If the downloaded data includes cell type labels, compare the clustering results against these labels. If the data lacks annotations, use marker genes to assign cell types and compare the results with published descriptions of the tissue or disease. The SingleR package provides automated cell type annotation based on reference datasets, and scUmaper offers marker-library-based annotation.
Common Failure Patterns in scRNA-seq Data Access
Incomplete Metadata Prevents Interpretation
Datasets with sparse metadata limit the conclusions that can be drawn. If the GEO record does not specify the tissue source, the disease state, or the experimental conditions, the data may still be analyzable but the biological interpretation becomes speculative. Researchers should document these limitations and avoid overinterpreting results from poorly annotated datasets.
File Format Mismatches Cause Loading Errors
Attempting to load a file in the wrong format is a common source of errors. For example, loading a raw feature-barcode matrix when the analysis expects a filtered matrix produces different cell counts and may include empty droplets. Loading a CSV file when the tool expects an H5 file generates a parsing error. Checking the file extension and reading the tool documentation before loading prevents these issues.
Version Incompatibilities Between Tools
Seurat, Scanpy, and Bioconductor packages update frequently, and version incompatibilities can break analysis scripts. A script written for Seurat v4 may not run in Seurat v5 without modification. The Bioconductor project manages package versions through its release system, which helps maintain compatibility within the R ecosystem. The nf-core documentation describes community pipeline standards that address version control and reproducibility, and the Galaxy Training Network provides accessible workflow training that covers tool version management.
Download Corruption and Incomplete Transfers
Large files downloaded over unstable connections may arrive corrupted or incomplete. The SRA Toolkit includes checksum verification, and GEO provides file sizes that can be compared against the downloaded file. Re-downloading files that fail verification prevents downstream analysis errors.
Batch Effects Across Multiple Datasets
Integrating data from multiple studies introduces batch effects, where technical differences between experiments obscure biological signals. Tools such as Harmony and scVI correct for batch effects, and the Celline workflow includes these tools in its integration pipeline. The MMIHCL framework uses hypergraph contrastive learning to integrate weakly linked multimodal data, which is relevant when combining datasets from different modalities or platforms.
Records and Documentation for Reproducible Data Access
Maintain a Data Access Log
Record the accession numbers, download dates, file versions, and processing steps for each dataset. This log supports reproducibility and provides the information needed for citations. The log should include:
- GEO or SRA accession number
- Dataset title and description
- Download date and source URL
- File names and checksums
- Software versions used for loading and processing
- Any format conversions or identifier mappings applied
Document Processing Decisions
Every processing decision affects the downstream analysis. Record the quality control thresholds, normalization method, and batch correction approach. The sctransform method, for example, uses regularized negative binomial regression for normalization and variance stabilization, and this choice affects variable gene selection, dimensional reduction, and differential expression. Documenting these choices allows other researchers to reproduce the analysis or understand its limitations.
Use Version Control for Analysis Scripts
Version control systems such as Git track changes to analysis scripts and provide a history of the analysis. The Carpentries lessons cover foundational computing skills including shell, Git, and programming, which support reproducible analysis practices. The nf-core documentation describes community pipeline standards that emphasize version control and reproducibility.
Limitations and Interpretation Boundaries
Processed Data Reflects Original Processing Choices
When using processed data from GEO, the analysis inherits the quality control and normalization decisions of the original authors. If the original analysis used lenient quality thresholds, the data may contain low-quality cells that affect downstream results. If the original analysis used aggressive filtering, rare cell populations may have been removed. Researchers should examine the quality metrics of the downloaded data and consider whether additional filtering is needed.
Dropout and Technical Noise Affect All scRNA-seq Data
Technical capture loss, or dropout, obscures true biological expression in scRNA-seq data. Existing imputation methods have difficulty distinguishing biological zeros from technical noise. The scZN framework addresses this by assuming that observed data arise from a combination of RNA's two-state transcription process and dropout, formulating imputation as nonnegative factorization. Across multiple real datasets, scZN captured true distributional characteristics at both gene and cell levels and suppressed spurious activation of genes that should not be expressed. Researchers should consider whether imputation is appropriate for their analysis and which method best suits their data.
Cell Type Annotation Is Inherently Interpretive
Cell type labels are assigned based on marker gene expression, which requires judgment and domain knowledge. Different annotation methods may produce different labels for the same clusters. The scUmaper framework provides marker-library-based annotation with lineage-marker incompatibility rules, which reduces the subjectivity of manual annotation. Researchers should validate annotations using multiple marker genes and compare results with published descriptions of the tissue.
Cross-Dataset Integration Has Statistical Limits
Integrating data from different studies, platforms, or modalities introduces challenges that no computational method fully resolves. Weakly linked modalities, where correlations between data types are low, present particular difficulties. The MMIHCL framework addresses these challenges using hypergraph contrastive learning, but the quality of integration depends on the underlying data. Researchers should evaluate integration quality using metrics such as the single-cell integration benchmark (scIB) score and should interpret cross-dataset results cautiously.
Safety and Regulatory Context for Data Use
Data Use Agreements and Licensing
Public repositories provide data for research use, but some datasets have restrictions on redistribution or commercial use. Check the GEO record for any data use restrictions before downloading. The NCBI data usage policies describe the terms under which data can be used and redistributed. Researchers should comply with these terms and cite the data sources in publications.
Privacy Considerations for Human Data
Human scRNA-seq data may contain sensitive information, even after de-identification. Some datasets include clinical metadata that could identify individuals if combined with other information. Researchers should handle human data according to their institutional review board requirements and applicable regulations. The NCBI provides guidance on human data submission and access, including controlled access for sensitive datasets.
Reproducibility Standards
Funding agencies and journals increasingly require data availability statements and reproducible analysis workflows. The Galaxy Training Network provides accessible workflow training that emphasizes reproducibility, and the nf-core documentation describes community pipeline standards. Following these standards supports compliance with publication requirements and facilitates scientific validation.
Professional Escalation Criteria
When to Seek Specialized Support
Some data access challenges require specialized expertise. Consider escalating to a bioinformatics core facility or collaborating with a computational biologist when:
- The dataset requires extensive preprocessing or custom alignment
- The analysis involves integrating multiple datasets with complex batch structures
- The data format is unfamiliar or the tool documentation is insufficient
- The computational requirements exceed available local resources
- The biological interpretation requires specialized domain knowledge
When to Question Data Quality
If the downloaded data fails quality checks, investigate before proceeding. Persistent issues may indicate problems with the dataset itself, and using flawed data produces unreliable results. Consider whether an alternative dataset better suits the analysis. Published studies provide examples of quality assessment, such as the COPD study that used the Seurat package for quality control, dimension reduction, and cell identification of scRNA-seq data.
When to Consult Repository Support
Repository support teams can help with access issues, file format questions, and metadata inquiries. The NCBI and EMBL-EBI provide help desks and documentation. If a dataset appears incomplete or the metadata is inconsistent, contacting the repository or the original authors may resolve the issue.
Building a Dataset Selection Scorecard for Single-Cell RNA-seq Projects
Researchers often discover after downloading a scRNA-seq dataset that it lacks critical metadata, uses an incompatible file format, or contains too few cells for the planned analysis. These failures waste days of work and delay the entire project. A structured selection scorecard applied before download prevents most of these problems. This section provides a decision framework that ranks candidate datasets against the specific requirements of the downstream analysis, along with a record system for tracking dataset evaluations and a troubleshooting method for common access failures.
Why a Scorecard Is Necessary for Single-Cell Data
Bulk RNA-seq datasets are relatively forgiving. A researcher can download a series matrix, load it into a differential expression tool, and obtain interpretable results even when the metadata is sparse. Single-cell datasets are different. The analysis pipeline includes quality control filtering, doublet removal, normalization, clustering, and cell type annotation, and each step depends on information that may or may not be present in the repository record. A dataset with 500 cells cannot support a rare cell population analysis that requires thousands of cells. A dataset deposited as raw FASTQ files only cannot be loaded into Seurat without substantial computational work. A dataset with no tissue annotation cannot be interpreted biologically.
The scorecard addresses these issues by forcing an explicit comparison between what the analysis requires and what the dataset provides. It converts an informal judgment into a documented decision that can be reviewed, shared with collaborators, and revisited if the analysis plan changes.
The Seven-Point Dataset Selection Scorecard
Apply the following seven criteria to every candidate dataset before downloading. Score each criterion as pass, caution, or fail, and record the evidence for each score in the dataset evaluation log described below.
Criterion 1: Cell Count Sufficiency
The number of cells in the dataset must support the planned analysis. A cell type identification study requires enough cells per expected population to achieve statistical power. A trajectory analysis requires sufficient cells across the developmental continuum. A rare population study requires a much larger total cell count than a study of abundant populations.
Check the GEO record for the total number of cells reported. Published studies that used the dataset provide reference numbers. For example, one endometrial carcinoma study detected 33,408 genes in 33,162 cells from scRNA-seq data, and a bladder cancer study identified 19 cell subpopulations and 7 core cell types from its dataset. Compare these numbers against the minimum required for the analysis. If the record does not state the cell count, examine the supplementary matrix file dimensions after download, but treat the absence of this information as a caution flag.
Criterion 2: Metadata Completeness for Biological Interpretation
The dataset must include the metadata needed to answer the biological question. At minimum, the record should specify the organism, tissue, disease state, and experimental conditions. For disease studies, the record should distinguish case from control samples. For developmental studies, the record should specify the time points or stages collected.
Datasets that lack this information may still be analyzable, but the interpretation becomes speculative. The GEO record summary and overall design sections usually contain this information. Published papers that generated the data provide additional context. If the metadata is insufficient to assign samples to experimental groups, the dataset fails this criterion.
Criterion 3: File Format Compatibility with Planned Tools
The supplementary files must be loadable by the analysis software. Seurat and Scanpy read 10x Genomics MEX and H5 formats directly. Other formats require conversion or custom parsing. The Bioconductor project provides packages and workflows for low-level analysis of scRNA-seq data, including quality control, normalization, clustering, and marker gene detection, and these packages have specific input format requirements.
Before downloading, list the files in the GEO supplementary section and confirm that at least one file matches the expected format. If the dataset provides only raw FASTQ files, the analysis requires an alignment step that may exceed the available computational resources. If the dataset provides only per-cell files from Smart-seq2, the loading process differs substantially from 10x data. The scorecard should record the planned tool and the confirmed file format.
Criterion 4: Raw Data Availability for Validation or Reprocessing
Processed data inherits the processing choices of the original authors, including their aligner, reference genome version, and quality thresholds. For most downstream analyses, processed data is sufficient. However, some situations require raw data. If the analysis plan includes custom alignment, alternative quantification, or re-processing with different quality thresholds, the dataset must include raw FASTQ files in SRA.
Check whether the GEO record links to SRA accessions. The SRA Toolkit provides the prefetch and fasterq-dump commands for downloading raw data. Raw data downloads are substantially larger than processed data, and a single scRNA-seq sample can require tens of gigabytes of storage. The scorecard should record whether raw data is available and whether the analysis plan requires it.
Criterion 5: Technical Quality Indicators
The dataset should include information about sequencing depth, platform, and quality metrics. The number of unique molecular identifiers per cell, the number of genes detected per cell, and the percentage of mitochondrial reads are standard quality indicators. The sctransform package offers a modeling framework for normalization and variance stabilization that removes the influence of technical characteristics while preserving biological heterogeneity, but this method requires UMI-based data.
If the GEO record or the associated publication reports these metrics, record them in the evaluation log. If the dataset lacks this information, the quality assessment must be performed after download, which adds time to the project. The absence of reported quality metrics is a caution flag, not an automatic failure.
Criterion 6: Batch Structure and Integration Requirements
Datasets generated across multiple batches, samples, or sequencing runs require batch correction before integration. The batch structure should be documented in the metadata. Tools such as Harmony and scVI correct for batch effects, and the Celline workflow includes these tools in its integration pipeline. The MMIHCL framework uses hypergraph contrastive learning to integrate weakly linked multimodal data, which is relevant when combining datasets from different modalities or platforms.
If the dataset includes samples from multiple patients or multiple experimental runs, the metadata must identify the batch for each cell. Without this information, batch correction cannot be applied, and the integration may produce spurious clusters driven by technical variation instead of biology.
Criterion 7: Data Use Restrictions
Some datasets have restrictions on redistribution or commercial use. The GEO record may include a data use agreement or a link to the original study's data availability statement. Human data may be subject to controlled access requirements. Check the record before downloading and record any restrictions in the evaluation log. The NCBI data usage policies describe the terms under which data can be used and redistributed.
The Dataset Evaluation Log
Maintain a structured log for every dataset considered for the project. This log serves as the decision record and supports reproducibility. Create one entry per candidate dataset with the following fields:
- Candidate identifier and GEO or SRA accession number
- Date of evaluation and evaluator name
- Biological question the dataset would address
- Score for each of the seven criteria with evidence
- Overall recommendation with justification
- Decision and date of decision
The log should be stored with the analysis scripts in version control. The Carpentries lessons cover foundational computing skills including shell, Git, and programming, which support reproducible analysis practices. The nf-core documentation describes community pipeline standards that emphasize version control and reproducibility.
Applying the Scorecard to a Realistic Scenario
Consider a researcher planning to study immune cell populations in gastric cancer. The analysis requires at least 10,000 cells, cell type annotation using marker genes, and comparison between tumor and adjacent normal tissue. The researcher identifies three candidate datasets from GEO.
The first candidate is GSE167297, used in a published gastric cancer study that combined transcriptomics and single-cell RNA sequencing. The study reported COL5A2 expression analysis in gastric cancer and used the dataset for scRNA-seq analysis. The GEO record includes processed count matrices and metadata describing tumor and normal samples. The cell count exceeds 10,000. The file format is compatible with Seurat. Raw data is available in SRA. The dataset passes all seven criteria.
The second candidate is a dataset from a related gastrointestinal study that provides only raw FASTQ files. The cell count is not stated in the record. The metadata describes the tissue but does not distinguish tumor from normal samples. The dataset fails the metadata completeness criterion and receives a caution for cell count sufficiency. The researcher records the evaluation and moves to the third candidate.
The third candidate is a dataset from a different cancer type. The metadata is complete and the file format is compatible, but the tissue does not match the research question. The dataset fails the biological relevance check that precedes the formal scorecard. The researcher records this decision and selects the first candidate.
Troubleshooting Method for Download and Loading Failures
When a download or loading step fails, work through the following troubleshooting sequence before contacting repository support.
Step 1: Verify File Integrity
Large files downloaded over unstable connections may arrive corrupted or incomplete. Compare the downloaded file size against the size listed in the GEO record. If the repository provides checksums, verify the file against the checksum. The SRA Toolkit includes checksum verification for raw data downloads. Re-download any file that fails verification.
Step 2: Confirm File Format Against Tool Documentation
Loading errors often result from format mismatches. Check the file extension and the first few lines of the file to confirm the format. A Matrix Market file begins with a header line that includes the dimensions. An H5 file is a binary format that cannot be inspected as text. Read the tool documentation to confirm the expected input format. The Bioconductor project provides package documentation that describes input requirements for each function.
Step 3: Check for Version Incompatibilities
Seurat, Scanpy, and Bioconductor packages update frequently, and version incompatibilities can break analysis scripts. A script written for Seurat v4 may not run in Seurat v5 without modification. Check the package version and the dataset format. The Galaxy Training Network provides accessible workflow training that covers tool version management, and the nf-core documentation describes community pipeline standards that address version control.
Step 4: Inspect the Data Structure
If the file loads but produces unexpected results, inspect the data structure. Check the matrix dimensions against the GEO record description. Verify that the gene identifiers match the reference genome used in the analysis. Some datasets use gene symbols, while others use Ensembl IDs or Entrez IDs. Conversion between identifier types may be necessary.
Step 5: Compare Against Published Quality Metrics
Published studies that used the dataset provide reference numbers for validation. For example, a COPD study used the Seurat package for quality control, dimension reduction, and cell identification of scRNA-seq data from GSE173896. A colorectal cancer study identified seven main cell subtypes by scRNA-seq analysis. If the downloaded data produces substantially different numbers, investigate the discrepancy before proceeding.
Step 6: Escalate to Repository Support
If the troubleshooting sequence does not resolve the issue, contact the repository support team. The NCBI and EMBL-EBI provide help desks and documentation. Include the accession number, the error message, and the steps already taken in the support request. If the dataset appears incomplete or the metadata is inconsistent, contacting the original authors may also resolve the issue.
Common Failure Patterns and Their Scorecard Prevention
The scorecard prevents the most common dataset access failures before they occur. A dataset with insufficient cell count fails the first criterion and is rejected before download. A dataset with incomplete metadata fails the second criterion. A dataset with incompatible file formats fails the third criterion. A dataset with no raw data fails the fourth criterion when raw data is required. A dataset with poor quality indicators fails the fifth criterion. A dataset with undocumented batch structure fails the sixth criterion. A dataset with use restrictions fails the seventh criterion.
The evaluation log provides the documentation needed to justify dataset selection in publications and to collaborators. It also provides a record that can be revisited if the analysis plan changes. A dataset rejected for insufficient cell count for one analysis may be suitable for another analysis with different requirements. The log preserves the evaluation and allows the researcher to revisit the decision without repeating the assessment.
Integration with the Broader Data Access Workflow
The scorecard fits between the search strategy and the download step in the overall workflow. After identifying candidate datasets through GEO searches or published accession numbers, apply the scorecard to each candidate. Select the highest-scoring dataset that meets all pass criteria. Then proceed to the download and quality control steps.
The scorecard also supports multi-dataset projects. When the analysis requires integrating data from multiple studies, apply the scorecard to each dataset and record the batch structure for each. The integration step requires knowing which cells come from which study and which batch. The evaluation log provides this information in a structured format.
Limitations of the Scorecard Approach
The scorecard cannot substitute for domain expertise. A researcher who does not know the expected cell types in a tissue cannot evaluate whether the dataset contains the relevant populations. A researcher who does not know the minimum cell count for a trajectory analysis cannot apply the first criterion effectively. The scorecard structures the decision but does not make it.
The scorecard also cannot detect problems that only appear after download. A dataset may pass all seven criteria and still contain doublets, ambient RNA contamination, or batch effects that require correction. The quality control steps described in the main workflow address these issues after download. The scorecard reduces the risk of selecting an unsuitable dataset but does not eliminate the need for post-download validation.
The scorecard should be updated as the analysis plan evolves. A dataset selected for cell type identification may later be used for trajectory analysis, which requires different quality characteristics. The evaluation log should note these changes and the rationale for them.
Frequently Asked Questions
What is the difference between GEO and SRA for single-cell data?
GEO stores processed and curated data, including count matrices and supplementary files, while SRA stores raw sequencing reads. For most downstream analyses, GEO processed data is sufficient and requires less computational effort. SRA data is needed when the analysis requires custom alignment or when the processed data is not available.
How do I download 10x Genomics data from GEO?
Look for the supplementary files in the GEO record, which typically include the filtered feature-barcode matrix in MEX or H5 format. Download these files and load them using the Read10X function in Seurat or the corresponding function in Scanpy. Verify that the files match the dataset description before proceeding with analysis.
What file formats are used for single-cell RNA-seq data?
Common formats include MEX (matrix, features, barcodes), H5, CSV, and TSV for processed data, and FASTQ for raw reads. The 10x Genomics platform produces MEX and H5 formats, while other platforms may produce different formats. Check the GEO record to confirm the format before downloading.
How do I convert raw SRA data to a count matrix?
Download the raw data using the SRA Toolkit, convert to FASTQ format, then run an alignment and quantification pipeline such as Cell Ranger for 10x data or STARsolo for other platforms. This process requires substantial computational resources and bioinformatics expertise.
What quality control steps should I perform after downloading scRNA-seq data?
Calculate the number of UMIs per cell, genes detected per cell, and percentage of mitochondrial reads. Filter low-quality cells and doublets using tools such as Scrublet or scUmaper. Check for ambient RNA contamination and apply correction if needed. Document all quality control decisions for reproducibility.
How do I choose between Seurat, Scanpy, and Bioconductor for analysis?
The choice depends on programming language preference and analysis goals. Seurat is an R package with extensive documentation and community support. Scanpy is a Python package that integrates with the broader Python ecosystem. Bioconductor provides a curated set of R packages for genomic analysis with a formal release system. All three support standard scRNA-seq workflows.
Can I integrate single-cell data from multiple studies?
Yes, but integration requires batch correction to account for technical differences between studies. Tools such as Harmony and scVI perform batch correction, and integrated tools such as Celline wrap these methods into automated workflows. Evaluate integration quality using metrics such as the scIB score and interpret cross-dataset results cautiously.
How do I cite single-cell datasets in my publications?
Cite the original study that generated the data and include the GEO or SRA accession numbers in the data availability statement. Follow the citation format required by the target journal and the repository guidelines. Proper citation ensures that data generators receive credit and supports scientific reproducibility.
Related Bioinformatics Guides
- RNA-Seq Databases: Accessing and Using Public RNA-Seq Data
- Single-Cell Sequencing Depth: How Much Is Enough?
- Single-Cell RNA Sequencing Quality Control: A Practical Guide to Filtering and Metrics
- Modern Transcriptomics: From Bulk RNA-Seq to Single-Cell and Spatial Resolution
- Single-Cell RNA Sequencing Depth: A Cost-Benefit Analysis for Experimental Design
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
- NCBI Data Resources. National Center for Biotechnology Information.
- EMBL-EBI Training. European Bioinformatics Institute.
- Bioconductor. Bioconductor Project.
- Galaxy Training Network. Galaxy Project.
- nf-core Documentation. nf-core.
- The Carpentries Lessons. The Carpentries.
- Integrating single-cell RNA-seq and spatial transcriptomics reveals MDK-NCL dependent immunosuppressive environment in endometrial carcinoma.. Frontiers in immunology, 2023.
- Comprehensive analysis of scRNA-Seq and bulk RNA-Seq reveals dynamic changes in the tumor immune microenvironment of bladder cancer and establishes a prognostic model.. Journal of translational medicine, 2023.
- Integrated analysis of single-cell RNA-seq and bulk RNA-seq to unravel the molecular mechanisms underlying the immune microenvironment in the development of intestinal-type gastric cancer.. Biochimica et biophysica acta. Molecular basis of disease, 2024.
- Identification of novel biomarkers related to neutrophilic inflammation in COPD.. Frontiers in immunology, 2024.
- Construction of an immune predictive model and identification of TRIP6 as a prognostic marker and therapeutic target of CRC by integration of single-cell and bulk RNA-seq data.. Cancer immunology, immunotherapy : CII, 2024.
- Identification of Five Hub Genes Based on Single-Cell RNA Sequencing Data and Network Pharmacology in Patients With Acute Myocardial Infarction.. Frontiers in public health, 2022.
- Integrative single-cell and bulk RNA-seq analyses identify CD4(+) T-cell subpopulation infiltration and biomarkers of regulatory T cells involved in mediating the progression of atherosclerotic plaque.. Frontiers in immunology, 2024.
- Machine learning-based prognostic model of lactylation-related genes for predicting prognosis and immune infiltration in patients with lung adenocarcinoma.. Cancer cell international, 2024.
- Prior-guided factorization for reliable imputation of scRNA-seq data.. 2026.
- Single-cell data integration across weakly linked modalities.. 2026.
- scUmaper: An automated framework for doublet removal and cell-type annotation in single-cell transcriptomics.. 2026.
- ScRNA-seq Data Reveal Gene Upregulation and Downregulation in Oxygen-Induced Retinopathy.. 2026.
- Celline: a flexible tool for one-step retrieval and integrative analysis of public single-cell RNA sequencing data.. 2025.
- Secondary Immune Surveillance via upregulation of C-type lectin receptors.. 2026.
- COL5A2 is a prognostic-related biomarker and correlated with immune infiltrates in gastric cancer based on transcriptomics and single-cell RNA sequencing. BMC Medical Genomics, 2023.
- Identification of the CD8+ T-cell exhaustion signature of hepatocellular carcinoma for the prediction of prognosis and immune microenvironment by integrated analysis of bulk- and single-cell RNA sequencing data. Translational Cancer Research, 2023.
- An integrative analysis of single-cell and bulk transcriptome and bidirectional mendelian randomization analysis identified C1Q as a novel stimulated risk gene for Atherosclerosis. Frontiers in Immunology, 2023.
- Normalization and variance stabilization of single-cell RNA-seq data using regularized negative binomial regression. Genome Biology, 2019.
- A step-by-step workflow for low-level analysis of single-cell RNA-seq data with Bioconductor. F1000Research, 2016.
- Characterization of cancer-related fibroblasts (CAF) in hepatocellular carcinoma and construction of CAF-based risk signature based on single-cell RNA-seq and bulk RNA-seq data. Frontiers in Immunology, 2022.
- Exploring the molecular mechanisms and shared gene signatures between celiac disease and ulcerative colitis based on bulk RNA and single-cell sequencing: Experimental verification. International Immunopharmacology, 2024.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.