Why Can't I Find the RNA-seq Data I Need? Troubleshooting Common Issues in Public Repositories

By Dr. Zubair Khalid, DVM, MS, PhD ·

Why Can't I Find the RNA-seq Data I Need? Troubleshooting Common Issues in Public Repositories

Key Takeaways

  • RNA-seq data resides in a layered system of repositories (e.g., NCBI SRA/GEO, EMBL-EBI ENA/ArrayExpress), with raw reads, processed data, and metadata often stored separately. Understanding these layers and their respective accession number formats (e.g., SRP for SRA studies, GSE for GEO series) is crucial for effective searching.
  • Incomplete or inconsistent metadata is a primary reason for search failures; researchers must employ multiple search terms, including synonyms and scientific/common names, and be prepared to examine raw files if experimental details are sparse.
  • Controlled access, particularly for human data due to privacy concerns, necessitates formal application processes via platforms like dbGaP, which can involve significant delays and require detailed research purpose justifications.
  • Downloading large RNA-seq datasets often requires specialized tools like the SRA Toolkit (e.g., fasterq-dump) and robust network connections; for very large files, cloud-based analysis platforms (e.g., Galaxy, nf-core) are recommended to avoid local storage and bandwidth limitations.
  • Data retrieval should be treated as a five-stage pipeline (Identification, Verification, Download, Validation, Usability), with a structured log to meticulously record each attempt, failure mode, and troubleshooting action, enabling systematic diagnosis of persistent access issues.
  • Reproducibility in RNA-seq analysis hinges on meticulous documentation of accession numbers, software versions, reference files, and download parameters, enabling others to replicate findings and facilitating troubleshooting of downstream analysis problems.

When you search for RNA-seq data in public repositories and come up empty, the problem is usually not that the data does not exist. More often, the issue lies in how the data was deposited, how you are searching, or how the repository organizes access. This article walks through the specific reasons why RNA-seq datasets can be difficult to locate or retrieve, and what you can do about each one. The focus is on practical steps you can take today, using the major public repositories and the tools they provide.

Understanding How RNA-seq Data Is Stored in Public Repositories

RNA-seq data lives in a layered system of databases, each with a different role. The raw sequencing reads, the processed count tables, and the metadata describing the experiment are often stored separately. Knowing which layer holds what you need is the first step to finding it.

The National Center for Biotechnology Information (NCBI) hosts the Sequence Read Archive (SRA), which stores raw sequencing data, and the Gene Expression Omnibus (GEO), which stores curated gene expression datasets including processed data and experimental metadata. The European Bioinformatics Institute (EMBL-EBI) runs equivalent resources, including the European Nucleotide Archive (ENA) and ArrayExpress. These are the primary places where RNA-seq data is deposited, and each has its own search interface and accession number system.

A common point of confusion is that a single study may have multiple accession numbers. The overall study might have a BioProject accession, the raw data might have an SRA accession, and the processed data might have a GEO accession. If you only have one of these numbers, you may need to search across multiple databases to find all the associated files.

The NCBI provides a unified search system that covers all of its databases at once. If you are unsure which database holds your data, start with the main NCBI search page and enter the accession there. The results will show you which database contains the record. This single step resolves a large fraction of "I cannot find the data" problems.

At a Glance: Common RNA-seq Data Retrieval Problems and Solutions

Problem You EncounterLikely CauseFirst Action to TryEscalation If It Fails
Search returns no results for a known accessionAccession is from a different database or has a typoVerify the accession prefix and search the correct databaseUse the NCBI cross-database search or contact the repository help desk
Metadata is missing or incompleteDepositor submitted minimal informationSearch for the associated BioProject or publicationContact the corresponding author of the associated paper
Raw data files are restricted or controlled accessHuman data with privacy protectionsCheck if you qualify for access and apply through the correct processContact the data access committee listed in the repository record
Download fails or times outLarge file sizes or network issuesUse the SRA Toolkit or a download managerTry a different download method or use a cloud-based analysis platform
Processed data is absentDepositor only submitted raw readsRun your own alignment and quantification pipelineCheck for supplementary files in the associated publication
Paired-end reads are mislabeledFile naming or metadata errorInspect the read files directly for adapter sequencesRe-run quality control and verify read pairing before analysis
Data is on an unfamiliar platformStudy used a less common repositorySearch the publication for the data availability statementUse cross-repository search tools or contact the authors

The Role of Accession Numbers in Finding RNA-seq Data

Accession numbers are the keys to public genomic data. Each database uses a specific format, and recognizing the format tells you where to search. NCBI SRA accessions typically start with SRR for runs, SRX for experiments, SRP for studies, and SRA for the archive itself. GEO accessions start with GSE for series, GSM for samples, and GPL for platforms. If you have a number that does not match the database you are searching, you will get no results even though the data exists.

When you have a publication but no accession number, check the data availability statement. Most journals now require authors to state where their data is deposited and to provide accession numbers. This statement is usually at the end of the paper, before the references. If the statement is missing or vague, the next step is to search for the authors' names or the study title in the repository directly.

The EMBL-EBI training materials cover how to search biological databases effectively, including strategies for dealing with inconsistent metadata. Their guidance emphasizes understanding how each database structures its records and using the advanced search features to combine terms.

Why Metadata Problems Hide RNA-seq Data from Searches

Metadata is the descriptive information attached to a dataset, including the organism, tissue type, condition, and sequencing platform. When metadata is incomplete or inconsistent, searches fail even though the data is present. This is one of the most common reasons why researchers cannot find RNA-seq data.

Depositors may use different terms for the same thing. One study might label a sample as "liver" while another uses "hepatic tissue." One might specify "Mus musculus" while another uses "mouse." Search engines match exact terms, so variations in vocabulary cause missed results. To work around this, try multiple search terms for the same concept, including scientific names, common names, and synonyms.

Another metadata issue is the level of detail provided. Some deposits include only the minimum required fields, such as organism and platform, without describing the experimental conditions. If you are searching for a specific treatment or time point, you will not find it because that information was never entered. In this case, the data may still be useful, but you will need to download the raw files and examine them to determine the experimental design.

Restricted Access: When RNA-seq Data Is Not Publicly Downloadable

Some RNA-seq data is not openly available for download. This applies most often to human data, where privacy concerns and consent agreements limit who can access the raw sequencing reads. The data is deposited in the repository, and the record is visible in searches, but the actual files are under controlled access.

If you find a dataset that requires approval to access, the process varies by repository and by the data access committee that oversees the dataset. For NCBI databases, controlled access data is often managed through dbGaP, the database of Genotypes and Phenotypes. You will need to apply for access, describing your research purpose and agreeing to data use restrictions.

The application process takes time, often weeks or months, so plan accordingly. If your project has a deadline, start the application early. If you are not sure whether you qualify for access, contact the data access committee listed in the repository record. They can tell you what information you need to provide and what criteria they use to evaluate applications.

For non-human data, restricted access is less common but still possible. Some studies have embargo periods during which the data is only available to collaborators. After the embargo expires, the data becomes public. If you find a record with an embargo date, check whether that date has passed.

Using the SRA Toolkit to Download RNA-seq Data

When you have identified the dataset you need and confirmed that it is publicly available, the next step is downloading the files. The SRA Toolkit, provided by NCBI, is the standard tool for this purpose. It includes commands for downloading, validating, and converting SRA files to standard formats like FASTQ.

The most common command is fastq-dump or its newer replacement fasterq-dump. These commands take an SRA accession as input and produce FASTQ files that you can use in your analysis pipeline. The fasterq-dump command is generally faster and uses less memory than the older tool.

A frequent problem is that downloads fail or time out, especially for large datasets. RNA-seq experiments can produce many gigabytes of data per sample, and a full study can be hundreds of gigabytes. If your download fails partway through, the SRA Toolkit provides options for resuming interrupted downloads. You can also use the prefetch command to download the SRA file first, then convert it locally.

For very large datasets, consider using a cloud-based approach. The Galaxy Training Network provides tutorials on accessing and analyzing public data using cloud resources, which can avoid the need to download massive files to your local machine. Similarly, the nf-core documentation describes how to run analysis pipelines on cloud infrastructure, where the data can be accessed directly from the repository.

Alternative Repositories for RNA-seq Data

If you cannot find data in NCBI or EMBL-EBI, the data may be in a different repository. Several specialized repositories host RNA-seq data, and some journals require deposition in specific places.

The NCBI hosts several databases beyond SRA and GEO, including the Sequence Read Archive for raw data and the BioProject database for study-level information. The EMBL-EBI hosts the European Nucleotide Archive and the BioStudies database, which can link to data stored elsewhere.

Some fields have their own repositories. For example, single-cell RNA-seq data is often deposited in the Human Cell Atlas data portal or in the Single Cell Expression Atlas at EMBL-EBI. If you are looking for single-cell data and cannot find it in the general repositories, try these specialized resources.

The scQuery web server, described in a Nature Communications paper, provides an automated pipeline for downloading and processing publicly available scRNA-seq datasets. It integrates data from over 500 studies and provides a search interface for comparing cell types across datasets. If you are working with single-cell data, this resource can help you locate datasets you might not find through standard repository searches.

Quality Control Issues That Make RNA-seq Data Unusable

Sometimes you can find the data, download it, and then discover that it has quality problems. Poor quality data can produce misleading results, so it is important to assess data quality before investing time in analysis.

The first quality check is to examine the raw reads with a tool like FastQC. This tool reports on read quality scores, GC content, adapter contamination, and other metrics. If the quality scores are low or the adapter content is high, the data may need trimming before analysis.

A protocol published in Current Protocols describes a streamlined RNA-seq analysis workflow that includes quality control steps. The protocol uses FastQC for quality assessment and Trimmomatic for trimming adapters and low-quality reads. It also covers alignment with BWA or HISAT2, read counting with Subread, and normalization to transcripts per million (TPM). This workflow is designed to be reproducible and to require minimal local hardware by using Google Colab for the heavy computation.

Paired-end data requires special attention. If the read pairs are not properly matched, alignment will fail or produce incorrect results. Check that the forward and reverse read files contain the same number of reads and that the read names match between the two files. The Current Protocols workflow includes specific guidance on handling paired-end reads during analysis.

Common Failure Patterns When Working with Public RNA-seq Data

Several failure patterns recur when researchers try to use public RNA-seq data. Recognizing these patterns can save you time and frustration.

The first pattern is searching in the wrong database. If you have a GEO accession and search for it in SRA, you will not find it. The reverse is also true. Use the NCBI cross-database search to locate records across all NCBI databases at once.

The second pattern is assuming that processed data exists. Many deposits include only raw sequencing reads. The depositor may have performed their own analysis and reported the results in a paper, but they did not upload the processed count tables. In this case, you must run your own alignment and quantification pipeline to get expression values.

The third pattern is encountering file format issues. SRA files are not directly usable in most analysis tools. You must convert them to FASTQ format first. Similarly, some deposits include BAM files that are sorted or indexed in ways that are incompatible with your tools. Check the file formats before you start and convert as needed.

The fourth pattern is version mismatch. Reference genomes and annotation files are updated regularly. If you use a different version than the depositor used, your results will not match theirs. Record the exact versions of all reference files you use and note them in your methods.

How to Verify That Downloaded RNA-seq Data Is Complete

After downloading RNA-seq data, verify that you have everything you need before starting analysis. Missing files or corrupted files will cause errors that can be difficult to trace.

Check the file sizes against the repository record. The SRA record lists the number of reads and the file size for each run. If your downloaded file is much smaller than expected, the download may have been incomplete. The SRA Toolkit includes a validation command that checks the integrity of downloaded files.

Verify that you have both read files for paired-end data. The repository record will indicate whether the data is single-end or paired-end. For paired-end data, you should have two FASTQ files per sample, one for forward reads and one for reverse reads. The read names in the two files should match.

Check the read length and number of reads against the repository record. If the numbers do not match, you may have downloaded the wrong file or the file may be corrupted. The Current Protocols workflow includes steps for verifying data integrity before proceeding with analysis.

Using Cloud Platforms to Avoid Download Problems

Downloading large RNA-seq datasets to a local computer can be impractical, especially if you have limited storage or bandwidth. Cloud platforms offer an alternative where you can access and analyze data without downloading it first.

Google Colab provides free access to computing resources with a web browser interface. The Current Protocols workflow uses Google Colab for the normalization and visualization steps of RNA-seq analysis. This approach avoids the need for powerful local hardware and makes the analysis accessible to researchers with limited resources.

The Galaxy platform provides a web-based interface for bioinformatics analysis, including access to public data repositories. The Galaxy Training Network offers tutorials on using Galaxy for RNA-seq analysis, including how to import data from public repositories. This can be a good option if you are not comfortable with command-line tools.

The nf-core project provides standardized analysis pipelines that can be run on cloud infrastructure. Their documentation describes how to configure and run pipelines on various cloud platforms. If you have a large dataset and access to cloud computing, this approach can handle the analysis efficiently.

The Role of Reproducibility in RNA-seq Data Retrieval

Reproducibility is a central concern in RNA-seq analysis. If you cannot reproduce the results reported in a paper, the data may be insufficiently documented or the analysis may have been performed incorrectly.

The Bioconductor project provides tools and workflows for reproducible genomic analysis. Their documentation emphasizes the importance of recording the exact versions of all software and reference files used in an analysis. This allows others to reproduce your results and allows you to reproduce your own results months later.

The Carpentries lessons cover foundational computing skills, including version control with Git and reproducible data analysis practices. These skills are essential for managing RNA-seq analysis projects, where small changes in software versions or parameters can produce different results.

When you use public RNA-seq data, document everything. Record the accession numbers, the download date, the software versions, and the parameters used. This documentation is essential for reproducibility and for troubleshooting if something goes wrong.

Interpreting Results from Public RNA-seq Data

Once you have successfully retrieved and analyzed public RNA-seq data, you need to interpret the results correctly. This requires understanding the limitations of the data and the analysis.

Differential expression analysis identifies genes that change between conditions. The pyDESeq2 tool, used in the Current Protocols workflow, implements statistical methods for this purpose. The results include log fold changes and adjusted p-values for each gene. Genes with large fold changes and small p-values are considered differentially expressed.

Functional enrichment analysis identifies biological pathways that are overrepresented among differentially expressed genes. The g:Profiler tool, also used in the Current Protocols workflow, performs this analysis. The results can help you understand the biological significance of your findings.

Interpretation requires caution. Public datasets may have batch effects, where technical variation obscures biological variation. They may have been collected under conditions that are not fully described in the metadata. They may have quality problems that were not detected in the original analysis. Always consider these limitations when interpreting results.

Single-Cell RNA-seq Data: Special Retrieval Considerations

Single-cell RNA-seq data presents additional challenges for retrieval and analysis. The data volumes are larger, the file formats are more complex, and the analysis requires specialized tools.

A protocol for identifying recirculating thymic regulatory T cells describes the use of scRNA-seq and TCR-seq data. The protocol includes steps for single-cell sequencing and downstream computational analysis. It emphasizes the importance of careful cell sorting and quality control at each step.

The scQuery web server provides a way to search and compare publicly available scRNA-seq datasets. It uses supervised neural networks to classify cells and can help you determine whether a dataset contains the cell types you are interested in. This can save time by allowing you to identify relevant datasets before downloading them.

A protocol for conducting scATAC-seq analysis describes the challenges of working with single-cell chromatin accessibility data. The analytical workflow is complex and presents significant challenges for researchers new to the field. The protocol uses a publicly available dataset as an example and describes steps for data pre-processing and downstream analysis.

Bacterial RNA-seq Data: Unique Challenges

Bacterial RNA-seq data has its own set of challenges. The analysis pipeline differs from eukaryotic RNA-seq in several ways, including the need for strand-specific information and the handling of ribosomal RNA contamination.

A protocol for integrated analysis of bacterial RNA-seq and ChIP-seq data describes a systematic approach to combining transcriptomic and protein-DNA interaction data. The protocol includes the software environment, download and installation methods, and the analytical process. It provides mini-test data that can be used to verify that the analysis pipeline is working correctly.

The protocol emphasizes the importance of data consolidation, providing scripts for merging multiple files rapidly. This is particularly useful for bacterial studies, which often involve many samples and conditions.

When searching for bacterial RNA-seq data, be aware that the metadata may be less complete than for eukaryotic data. Bacterial studies often focus on specific strains or growth conditions, and this information may not be fully captured in the repository record.

Plant RNA-seq Data: Specialized Repositories and Formats

Plant RNA-seq data may be found in general repositories or in plant-specific resources. The data formats are generally the same as for other organisms, but the metadata may include plant-specific information such as cultivar, growth conditions, and tissue type.

A protocol for capturing the RNA-binding proteome from plants describes a method for isolating RNA-binding proteins crosslinked to RNA. While this is not strictly an RNA-seq protocol, it illustrates the diversity of RNA-related data that can be deposited in public repositories.

A scalable approach to evaluating plant microRNA trimming and tailing from small RNA-seq data is described in a Methods in Molecular Biology chapter. This protocol addresses the specific challenges of working with small RNA data, which requires different analysis approaches than standard mRNA-seq data.

When searching for plant RNA-seq data, try multiple search terms including the scientific name, common name, and cultivar. Plant researchers may use different naming conventions than researchers in other fields.

Nanopore Direct RNA-seq Data: Emerging Formats

Nanopore direct RNA-seq is an emerging technology that produces data in different formats than traditional short-read sequencing. The data files are larger and the analysis requires specialized tools.

A protocol for mapping 2'-O-methylation using nanopore direct RNA-seq data describes the use of the NanoNm tool for this purpose. The protocol includes steps for software installation, data collection, and training a machine learning model. It then details procedures for mapping 2'-O-methylation in both rRNA and mRNA in yeast and human cells.

When searching for nanopore RNA-seq data, be aware that the data may be deposited in different repositories than traditional RNA-seq data. The Oxford Nanopore Technologies community has its own data sharing platforms, and some data may only be available through these channels.

How to Contact Authors for Missing RNA-seq Data

When you cannot find RNA-seq data through any repository search, the final option is to contact the authors of the associated publication. This is often successful, as authors may have data that was not deposited or may be able to provide additional metadata.

Find the corresponding author's email address from the publication. Most journals list this information on the first page of the article. Write a concise email explaining what data you need, why you need it, and what you have already tried. Include the publication title and DOI, and any accession numbers you have found.

Be patient. Authors may take weeks to respond, and they may not have the data readily available. If you do not receive a response, try contacting a co-author or the institutional repository of the authors' institution.

When you receive the data, verify its integrity before using it. Check that the file formats are as expected and that the data matches the description in the publication. If the data was not deposited in a public repository, consider asking the authors to deposit it so that others can access it.

Professional Escalation Criteria for RNA-seq Data Problems

Some RNA-seq data problems cannot be solved by individual effort. When you have exhausted the standard troubleshooting steps, it may be time to escalate the issue.

Contact the repository help desk if you believe a record is incorrect or if you cannot access data that should be public. NCBI and EMBL-EBI both provide help desks that can assist with data access issues. Include the accession numbers and a description of the problem you are experiencing.

Contact the journal if the data availability statement in a publication is inaccurate or if the data is not available despite the statement saying it should be. Many journals now have policies requiring data availability, and they may be able to intervene.

Contact your institution's data librarian or bioinformatics support team. They may have experience with specific repositories or tools and can provide guidance on difficult cases.

If you are working with human data and believe you qualify for access but your application was denied, you can appeal the decision. The data access committee should provide information on the appeals process.

Records and Documentation for RNA-seq Data Retrieval

Keep detailed records of your data retrieval process. This documentation is valuable for reproducibility and for troubleshooting if problems arise later.

Record the following information for each dataset you retrieve:

  • The accession numbers for the study, experiment, and run
  • The date you downloaded the data
  • The tool and version used for downloading
  • The file sizes and checksums of downloaded files
  • Any problems encountered and how they were resolved

This information should be included in your methods section when you publish your results. It allows others to reproduce your data retrieval and analysis.

The Carpentries lessons emphasize the importance of documentation in scientific computing. Their guidance on project organization and version control is directly applicable to RNA-seq data retrieval and analysis.

Limitations of Public RNA-seq Data

Public RNA-seq data has inherent limitations that you should understand before using it. These limitations affect what conclusions you can draw from the data.

The metadata may be incomplete or inaccurate. Depositors may have made errors in describing their experimental conditions, or they may have omitted information that is important for interpretation. You cannot always trust that the metadata accurately describes the data.

The data quality may be variable. Some datasets have poor sequencing quality, high adapter contamination, or other problems. You should assess data quality before using it, and you should be prepared to discard datasets that do not meet your standards.

The data may not be representative. Public datasets are biased toward certain organisms, tissues, and conditions. If you are studying a rare condition or an unusual organism, you may find few or no public datasets that are relevant.

The analysis may not be reproducible. Even with good documentation, you may not be able to reproduce the exact results reported in a publication. Differences in software versions, parameters, and reference files can all affect the results.

Safety and Ethical Considerations for RNA-seq Data Use

Using public RNA-seq data carries ethical responsibilities, particularly for human data. Even when data is publicly available, you should consider the privacy implications for the individuals whose data you are using.

Human RNA-seq data may contain information that could identify individuals, even after de-identification. Be careful not to attempt to re-identify individuals from public data. This is both unethical and, in many jurisdictions, illegal.

When you publish results based on public data, acknowledge the original data producers. Cite the publication associated with the dataset and include the accession numbers. This gives credit to the researchers who generated the data and helps others find it.

If you use controlled access data, comply with all data use restrictions. These restrictions may limit what analyses you can perform and whether you can share the data with others. Violating these restrictions can have serious consequences.

Building a Structured Data Retrieval Log to Diagnose Persistent RNA-seq Access Failures

When you repeatedly fail to locate or download RNA-seq data, the problem is often not the repository but the absence of a systematic record of what you have tried. Researchers who keep no retrieval log tend to repeat the same failed searches, forget which accessions they have already verified, and lose track of which authors or help desks they have contacted. A structured retrieval log turns a frustrating trial-and-error process into a repeatable diagnostic procedure. This section provides a practical framework for building such a log, using it to identify the exact point of failure, and knowing when the problem lies with the data instead of with your search strategy.

The Five-Stage Retrieval Pipeline

Treat RNA-seq data retrieval as a pipeline with five distinct stages. Each stage has its own failure modes, and each requires different troubleshooting actions. By recording where in this pipeline your attempt fails, you can identify the specific cause and apply the correct remedy.

Stage 1: Identification. You have a publication, a grant number, or a topic and need to find the relevant dataset. Failure at this stage means you cannot locate any accession number. The cause is usually incomplete metadata, a missing data availability statement, or a search term mismatch. The remedy is to broaden your search terms, check the publication supplementary materials, or contact the corresponding author.

Stage 2: Verification. You have an accession number and need to confirm that the record exists and is publicly accessible. Failure here means the accession returns no results or shows restricted status. The cause is usually a typo, a wrong database, or controlled access. The remedy is to verify the accession prefix, use the NCBI cross-database search, or check the data access requirements.

Stage 3: Download. You have confirmed the record exists and are transferring the files. Failure here means the download times out, stops partway, or produces corrupted files. The cause is usually network instability, insufficient storage, or repository server load. The remedy is to use the SRA Toolkit prefetch command, resume interrupted downloads, or switch to a cloud-based access method.

Stage 4: Validation. You have downloaded the files and need to confirm they are complete and uncorrupted. Failure here means file sizes do not match the repository record, read counts are wrong, or paired-end files are mismatched. The cause is usually an interrupted download or a repository error. The remedy is to run the SRA Toolkit validation command and compare file properties against the record.

Stage 5: Usability. You have valid files but cannot use them in your analysis pipeline. Failure here means format incompatibility, missing index files, or reference version mismatches. The cause is usually a need for format conversion or additional processing. The remedy is to convert SRA to FASTQ, build or obtain the correct reference indexes, and document all software versions.

Designing the Retrieval Log

Create a simple table with one row per retrieval attempt. The columns should capture the information you need to diagnose failures and to document your methods for publication. The NCBI provides official descriptions of its databases and search systems, and the EMBL-EBI training materials cover effective search strategies, but neither will remember what you tried. Your log is the only record that captures your specific process.

Use these columns for each attempt:

  • Date of attempt. Record the exact date. Repository interfaces change, and data availability can shift when embargoes expire.
  • Dataset identifier. Record the full accession, including the prefix. Note whether it is a study, experiment, run, or series accession.
  • Source. Record where you found the accession, such as the publication data availability statement, a citation, or a colleague.
  • Repository searched. Record which database you searched, such as NCBI SRA, GEO, ENA, or ArrayExpress.
  • Search terms used. Record the exact terms, including synonyms and alternative spellings.
  • Stage reached. Record which of the five stages you reached before failure.
  • Failure description. Describe the specific error or problem in concrete terms.
  • Action taken. Record what you did in response, such as trying a different database or contacting the help desk.
  • Outcome. Record whether the action resolved the problem or whether you need to escalate.

Keep this log in a spreadsheet or a plain text file. The Carpentries lessons on project organization and documentation provide useful guidance on maintaining reproducible research records. A plain text file with comma-separated values is sufficient and has the advantage of being readable by any tool.

Using the Log to Diagnose Failure Patterns

After you have recorded several attempts, review the log for patterns. The pattern of failures tells you where the systemic problem lies.

If you consistently fail at Stage 1, the problem is likely with your search strategy or the completeness of the publication metadata. Try different search terms, including the organism scientific name, the tissue type, and the condition. Check whether the publication has a data availability statement and whether it provides accession numbers. If the statement is missing, search for the authors' names in the repository directly.

If you consistently fail at Stage 2, the problem is likely with the accession itself. Verify the prefix against the known formats for each database. A GEO series accession starting with GSE will not work in the SRA search. Use the NCBI cross-database search to locate records across all NCBI databases at once. If the record shows restricted status, determine whether you qualify for access and what application process applies.

If you consistently fail at Stage 3, the problem is likely with your network connection, storage capacity, or the repository server. Try the SRA Toolkit prefetch command to download the SRA file first, then convert it locally with fasterq-dump. If downloads continue to fail, consider using a cloud platform where the data can be accessed without a full download. The Galaxy Training Network provides tutorials on accessing public data through cloud resources, and the nf-core documentation describes running pipelines on cloud infrastructure.

If you consistently fail at Stage 4, the problem is likely with the download process or the repository record itself. Compare the downloaded file size and read count against the repository record. Run the SRA Toolkit validation command to check file integrity. For paired-end data, verify that the forward and reverse files contain the same number of reads and that read names match between files.

If you consistently fail at Stage 5, the problem is likely with your analysis environment instead of the data. Check that you have converted SRA files to FASTQ format. Verify that your reference genome and annotation versions match what you intend to use. The Bioconductor project provides documentation on reproducible genomic analysis, including the importance of recording software versions.

When the Log Points to a Data Problem

Sometimes the log reveals that the problem is not your process but the data itself. This happens when you have verified the accession, completed the download, validated the files, and still cannot use the data.

The most common data-level problem is missing processed data. Many deposits include only raw sequencing reads. The depositor performed their own analysis and reported results in a publication, but they did not upload the processed count tables. In this case, you must run your own alignment and quantification pipeline. A protocol published in Current Protocols describes a streamlined RNA-seq workflow using free tools including SRA Toolkit, FastQC, Trimmomatic, BWA or HISAT2, Samtools, and Subread, with normalization and visualization performed in Google Colab.

Another data-level problem is mislabeled paired-end reads. If the forward and reverse read files are swapped or contain mismatched read names, alignment will fail. The Current Protocols workflow includes specific guidance on handling paired-end reads during analysis.

A third data-level problem is poor sequencing quality. If the raw reads have low quality scores or high adapter contamination, the data may not be usable even after trimming. Run FastQC on a sample of the reads before committing to a full analysis. If the quality is unacceptable, consider whether the dataset is worth pursuing or whether you should look for an alternative.

Escalation Criteria Based on Log Evidence

Your retrieval log provides the evidence you need to escalate a problem to the repository help desk, the journal, or the data access committee. Without a log, your escalation request is vague and difficult to act on. With a log, you can show exactly what you tried, when you tried it, and where the process failed.

Contact the repository help desk when your log shows that a record is incorrect or that publicly accessible data cannot be downloaded. Include the accession numbers, the dates of your attempts, and a description of the errors you encountered. NCBI and EMBL-EBI both provide help desks that can investigate access issues.

Contact the journal when your log shows that the data availability statement in a publication is inaccurate. Many journals now have policies requiring data availability, and they may be able to intervene with the authors. Include the publication DOI and your retrieval log as evidence.

Contact the data access committee when your log shows that you have applied for controlled access and your application was denied or has not been processed. The committee should provide information on the appeals process. Your log demonstrates that you have followed the correct procedures and that the delay is not due to your own inaction.

Integrating the Log into Your Research Workflow

The retrieval log is not a one-time exercise. It should be part of your standard workflow for every public dataset you use. When you publish results based on public data, include the relevant log entries in your methods section. This documentation allows others to reproduce your data retrieval and analysis, and it protects you if questions arise about the provenance of your data.

The EMBL-EBI training materials emphasize the importance of understanding how each database structures its records. Your log complements this knowledge by capturing the practical details of your specific retrieval attempts. Together, the training and the log give you a complete picture of how to find and use public RNA-seq data.

The nf-core documentation describes community standards for reproducible workflows, including the importance of recording data sources and versions. Your retrieval log is the data-source component of this standard. When you run an nf-core pipeline, you can reference the log entries that document where each input file came from and how it was obtained.

A protocol for integrated analysis of bacterial RNA-seq and ChIP-seq data provides an example of how detailed documentation supports reproducibility. The protocol includes the software environment, download and installation methods, and the analytical process, along with mini-test data for verification. Your retrieval log serves a similar purpose for the data acquisition stage of your project.

Frequently Asked Questions

Why does my search return no results when I know the data exists?

The most common cause is searching in the wrong database. Each repository has its own accession number format, and a number from one database will not work in another. Use the NCBI cross-database search to locate records across all databases at once. If you have a publication, check the data availability statement for the correct accession numbers.

What should I do if the metadata for a dataset is incomplete?

Incomplete metadata is common in public repositories. Try searching with different terms, including synonyms and alternative names for the organism or condition. If you still cannot find what you need, download the raw data and examine it directly to determine the experimental design. You can also contact the authors for additional information.

How do I access controlled access RNA-seq data?

Controlled access data requires an application to the data access committee. The application process varies by repository and dataset. You will need to describe your research purpose and agree to data use restrictions. Start the application early, as the process can take weeks or months.

Why does my download keep failing or timing out?

Large RNA-seq datasets can be difficult to download over unstable connections. Use the SRA Toolkit's prefetch command to download the SRA file first, then convert it locally. You can also resume interrupted downloads. For very large datasets, consider using a cloud platform where the data can be accessed without downloading.

What do I do if the processed data is not available?

Many deposits include only raw sequencing reads. You will need to run your own alignment and quantification pipeline to get expression values. The Current Protocols workflow provides a step-by-step approach that uses free tools and can be run on Google Colab.

How can I tell if downloaded data is corrupted or incomplete?

Compare the file sizes and read counts against the repository record. The SRA Toolkit includes a validation command that checks file integrity. For paired-end data, verify that the forward and reverse read files contain the same number of reads and that read names match.

Where can I find single-cell RNA-seq data?

Single-cell data is often deposited in specialized repositories such as the Single Cell Expression Atlas or the Human Cell Atlas data portal. The scQuery web server provides a search interface for publicly available scRNA-seq datasets and can help you identify relevant data.

What should I do if I cannot find the data anywhere?

Contact the corresponding author of the associated publication. Explain what data you need and what you have already tried. If the author does not respond, try contacting co-authors or the institutional repository. You can also contact the journal if the data availability statement is inaccurate.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.