How to Use SRA Toolkit to Download RNA-seq Raw Data: A Step-by-Step Tutorial

By Dr. Zubair Khalid, DVM, MS, PhD ·

How to Use SRA Toolkit to Download RNA-seq Raw Data: A Step-by-Step Tutorial

Key Takeaways

  • The SRA Toolkit, comprising prefetch and fasterq-dump, is the primary NCBI-provided command-line interface for downloading raw RNA-seq FASTQ files from the Sequence Read Archive (SRA).
  • prefetch downloads compressed SRA files to a local cache, supporting resumable transfers and optional Aspera protocol for accelerated downloads, while fasterq-dump converts these SRA files into standard FASTQ format, with --gzip for compression and --threads for parallel processing.
  • Data integrity verification is critical; vdb-validate checks SRA file structure and checksums, and wc -l or grep -c "^@" on FASTQ files confirms read counts, especially for paired-end data where both mates must have identical read numbers.
  • Establishing a structured project directory (e.g., raw_sra/, fastq/) and adhering to consistent file naming conventions (e.g., using SRA run accessions) are essential for managing large RNA-seq datasets and ensuring reproducibility.
  • Before downstream analysis, quality control using tools like FastQC is mandatory to assess per-base quality scores, GC content, and identify adapter contamination, which may necessitate trimming with tools like Trimmomatic.
  • Troubleshooting common issues involves verifying network connectivity, ensuring sufficient disk space, checking file permissions, and understanding that configuration errors can often be resolved by resetting with vdb-config --restore.

RNA sequencing has become a standard method for studying gene expression across entire genomes, but the computational steps required to obtain and process raw data often create a barrier for researchers who lack advanced command-line experience. This tutorial addresses the specific problem of downloading raw RNA-seq FASTQ files from the NCBI Sequence Read Archive (SRA) using the SRA Toolkit. You will learn how to install the toolkit, locate accession numbers for your dataset of interest, use prefetch and fasterq-dump to retrieve files, verify data integrity, and troubleshoot common issues that arise during the download process. The instructions assume you have basic familiarity with a Unix-like terminal environment, whether on a local Linux machine, macOS, or Windows Subsystem for Linux.

Understanding the NCBI Sequence Read Archive and Its Role in RNA-seq Research

The National Center for Biotechnology Information (NCBI) maintains the Sequence Read Archive as a public repository for high-throughput sequencing data, including the raw output from RNA-seq experiments [<a href="#ref-1">1</a>]. When researchers publish a study involving transcriptome analysis, they are expected to deposit the raw sequencing reads in SRA so that other investigators can reproduce the analysis, reanalyze the data with different methods, or combine datasets for meta-analyses. The SRA database stores data in a compressed format that preserves the original sequencing reads along with quality scores, and it provides several tools for retrieving these files.

For RNA-seq research, the SRA serves as the primary source of raw data for a wide range of downstream applications. Published protocols describe workflows that begin with downloading raw data from NCBI GEO or SRA, then proceed through quality checking, adapter trimming, alignment to reference genomes, read counting, normalization, and differential expression analysis [<a href="#ref-2">2</a>]. Automated pipelines such as Octopus-toolkit have been developed to retrieve and process large sets of next-generation sequencing data from the Gene Expression Omnibus repository, including RNA-seq datasets from human, mouse, plant, zebrafish, fruit fly, worm, and yeast genomes [<a href="#ref-3">3</a>]. These pipelines rely on the SRA Toolkit as a core component for downloading original files.

The practical implication for your research is straightforward. Before you can perform any RNA-seq analysis, whether you plan to use a local pipeline, a cloud-based platform, or a web service, you must first obtain the raw FASTQ files from SRA. The SRA Toolkit is the official software package provided by NCBI for this purpose, and mastering its basic commands will give you control over which files you download, where you store them, and how you verify their integrity.

Installing SRA Toolkit on Your System

The SRA Toolkit is distributed as precompiled binaries for Linux, macOS, and Windows operating systems. The installation process varies slightly depending on your platform, but the general approach involves downloading the appropriate archive, extracting it to a directory of your choice, and adding the binaries to your system path.

Linux Installation

For most Linux distributions, you will download a tar archive from the NCBI SRA Toolkit download page. After extracting the archive, you will find a directory containing executable files such as prefetch, fasterq-dump, and fastq-dump. To make these commands available from any directory, add the toolkit's bin directory to your PATH environment variable. This can be done by editing your shell configuration file, typically ~/.bashrc or ~/.profile, and adding a line such as export PATH=/path/to/sratoolkit/bin:$PATH.

macOS Installation

macOS users can follow a similar process, downloading the macOS build of the SRA Toolkit and adding it to their path. Alternatively, the toolkit is available through package managers such as Homebrew, which simplifies installation and updates. If you use Homebrew, the command brew install sratoolkit will install the toolkit and handle the path configuration automatically.

Windows Installation

Windows users have two primary options. The first is to download the Windows binaries and run them from the Command Prompt or PowerShell, adding the directory to the system PATH through the environment variables settings. The second option, which is often recommended for bioinformatics work, is to use Windows Subsystem for Linux (WSL) to run a Linux environment on your Windows machine. Published protocols for RNA-seq analysis have demonstrated the use of WSL to run command-line tools including the SRA Toolkit on Windows systems [<a href="#ref-2">2</a>]. This approach gives you access to the full range of Unix command-line utilities that are commonly used in bioinformatics workflows.

Verifying the Installation

After installation, verify that the toolkit is working correctly by running the command vdb-config --version or simply prefetch --version. These commands should display version information for the toolkit. The first time you run certain SRA Toolkit commands, the software may prompt you to configure your settings, including the location of your download cache and whether you want to enable cloud access. You can accept the default settings for most purposes, but you should be aware that the configuration file controls where temporary files are stored during downloads.

Locating RNA-seq Data in NCBI Databases

Before you can download RNA-seq data, you must identify the specific dataset you need and obtain its accession number. NCBI provides several search interfaces for this purpose, including the SRA database itself and the Gene Expression Omnibus (GEO) repository [<a href="#ref-1">1</a>]. The choice of search interface depends on how much information you have about the dataset.

Searching the SRA Database Directly

The SRA database can be searched using keywords, organism names, study accession numbers, or sample accession numbers. If you know the BioProject accession (for example, PRJNA907231) or the SRA study accession (for example, SRP followed by digits), you can search for these directly. The search results will show you the individual runs associated with the study, each with its own accession number beginning with SRR, ERR, or DRR depending on the sequencing center.

Searching GEO for Published Datasets

Many RNA-seq datasets are described in GEO records that include experimental details, sample information, and links to the underlying SRA data. A GEO series accession (beginning with GSE) will typically contain links to the SRA runs for each sample. When you find a GEO record of interest, look for the SRA links or the supplementary files section to identify the corresponding SRA accessions. Published RNA-seq datasets often include this information in their data availability statements, as demonstrated by studies that deposit raw data files at the NCBI SRA with specific accession numbers [<a href="#ref-4">4</a>].

Understanding Accession Number Formats

SRA accession numbers follow predictable patterns that help you identify the type of record. BioProject accessions begin with PRJNA, PRJEB, or PRJDB. SRA study accessions begin with SRP, ERP, or DRP. Sample accessions begin with SRS, ERS, or DRS. Run accessions, which correspond to individual sequencing runs, begin with SRR, ERR, or DRR. When you download data, you will typically work with run accessions, although the prefetch command can also accept study or sample accessions and download all associated runs.

Using prefetch to Download SRA Files

The prefetch command is the first step in the SRA Toolkit download workflow. It downloads SRA files from NCBI servers to your local machine, storing them in a compressed format that preserves the original data structure. The command is designed to be efficient, using checksums to verify data integrity and supporting resumable downloads if the connection is interrupted.

Basic prefetch Usage

The basic syntax for prefetch is straightforward. You provide one or more accession numbers, and the tool downloads the corresponding SRA files to your local cache directory. For example, to download a single run, you would use:

prefetch SRR12345678

To download multiple runs in a single command, list them separated by spaces:

prefetch SRR12345678 SRR12345679 SRR12345680

You can also provide a text file containing a list of accessions, one per line, using the --option-file flag. This approach is useful when you need to download a large number of runs from a study.

Specifying the Output Directory

By default, prefetch stores downloaded files in a cache directory defined in your configuration. You can override this with the --output-directory flag to specify where you want the files to be stored. For example:

prefetch SRR12345678 --output-directory /path/to/your/data

This creates a subdirectory named after the accession within your specified output directory, containing the SRA file.

Downloading Entire Studies

If you want to download all runs from a study, you can provide the study accession to prefetch. For example:

prefetch SRP123456

This command will download all runs associated with the study. This approach is convenient when you need the complete dataset for a study, but it can consume significant disk space and bandwidth for large studies. Check the study size before downloading to ensure you have adequate storage.

Using Aspera for Faster Downloads

The SRA Toolkit supports the Aspera transfer protocol, which can achieve faster download speeds than standard HTTP transfers. To use Aspera, you need the ascp binary installed on your system and must specify the --transport aspera flag. The vdb-config tool includes settings for Aspera configuration. Published pipelines have incorporated Aspera as an option for retrieving NGS data [<a href="#ref-3">3</a>]. However, Aspera requires a licensed client, and the configuration can be more complex than standard downloads. For most users, the default HTTPS transfer is sufficient, especially for smaller datasets.

Converting SRA Files to FASTQ with fasterq-dump

The SRA files downloaded by prefetch are not directly usable by most downstream analysis tools. You must convert them to FASTQ format, which is the standard input format for quality control tools, aligners, and quantification software. The fasterq-dump command performs this conversion, and it is the recommended tool for RNA-seq data because it is optimized for speed and produces output suitable for high-throughput analysis.

Basic fasterq-dump Usage

The basic syntax for fasterq-dump is:

fasterq-dump SRR12345678

This command reads the SRA file from your local cache and produces a FASTQ file (or two files for paired-end data) in your current working directory. For paired-end data, the output files are named with _1.fastq and _2.fastq suffixes.

Specifying Input and Output Locations

If your SRA files are stored in a non-default location, you can specify the path to the SRA file directly:

fasterq-dump /path/to/data/SRR12345678/SRR12345678.sra

To control where the FASTQ files are written, use the --outdir flag:

fasterq-dump SRR12345678 --outdir /path/to/fastq/output

Handling Paired-End Data

RNA-seq experiments frequently use paired-end sequencing, where each fragment is sequenced from both ends. The fasterq-dump command automatically detects whether the data is paired-end and produces two FASTQ files. You should verify that the read counts in the two files match, as this is an important quality check before proceeding with alignment. The --split-files flag is the default behavior for paired-end data and produces separate files for each mate.

Compressing FASTQ Output

FASTQ files are large, and compressing them can save significant disk space. The fasterq-dump command supports the --gzip flag to produce gzip-compressed output:

fasterq-dump SRR12345678 --gzip

This produces files with .fastq.gz extensions, which are accepted by most downstream tools. Compression adds some processing time but is generally worth the disk space savings for RNA-seq datasets.

Using fastq-dump as an Alternative

The older fastq-dump command is still available in the SRA Toolkit and can be used as an alternative to fasterq-dump. The fastq-dump command offers some additional options, such as the ability to split files into multiple parts, but it is generally slower than fasterq-dump. For most RNA-seq applications, fasterq-dump is the preferred choice due to its speed and efficiency.

Verifying Downloaded Data Integrity

After downloading and converting your data, you should verify that the files are complete and uncorrupted. The SRA Toolkit provides tools for this purpose, and there are additional checks you can perform to ensure your data is ready for downstream analysis.

Checking SRA File Integrity

The vdb-validate command checks the integrity of SRA files by verifying checksums and structural consistency:

vdb-validate SRR12345678

This command reports whether the file is valid and provides information about the number of reads and the read length. Run this command on your downloaded SRA files before converting them to FASTQ to catch any download errors early.

Verifying FASTQ File Completeness

After conversion, you can check the FASTQ files using standard Unix commands. The wc -l command counts the number of lines in a FASTQ file. Because each read occupies four lines in a FASTQ file, the number of reads is the line count divided by four. For paired-end data, the read counts in the two files should be identical. You can also use the grep command to count the number of read identifiers:

grep -c "^@" SRR12345678_1.fastq

This count should match the number of reads reported by vdb-validate.

Using md5sum for Checksum Verification

If you downloaded files from a source that provides checksums, you can use the md5sum command to verify that your downloaded files match the expected checksums. This is particularly important for large datasets where transmission errors are more likely. Some repositories provide checksum files alongside the data, and you should compare these against your downloaded files.

Organizing Your RNA-seq Data Directory Structure

A well-organized directory structure is essential for managing RNA-seq data, especially when working with multiple samples or studies. A consistent structure makes it easier to track which files correspond to which samples, simplifies the use of automated pipelines, and reduces the risk of errors in downstream analysis.

Recommended Directory Layout

A common approach is to create a project directory that contains subdirectories for each stage of the analysis. For example:

project/
├── raw_sra/
├── fastq/
├── qc/
├── trimmed/
├── alignment/
└── counts/

The raw_sra directory stores the SRA files downloaded by prefetch. The fastq directory contains the FASTQ files produced by fasterq-dump. Subsequent directories hold the outputs of quality control, trimming, alignment, and read counting steps. This structure keeps raw data separate from processed data, which is important for reproducibility.

Naming Conventions for Sample Files

Use consistent naming conventions for your sample files. A common convention is to use the SRA run accession as the base name, with suffixes indicating the processing stage. For example, SRR12345678_1.fastq.gz and SRR12345678_2.fastq.gz for raw paired-end reads, and SRR12345678_trimmed_1.fastq.gz for trimmed reads. If you are working with samples from your own experiment, you may prefer to use your own sample identifiers, but you should maintain a mapping file that links your identifiers to the SRA accessions.

Recording Metadata

Maintain a metadata file that records important information about each sample, including the SRA accession, the study accession, the organism, the tissue or cell type, the experimental condition, and the sequencing platform. This metadata is essential for interpreting your analysis results and for sharing your data with collaborators or reviewers. Published RNA-seq datasets typically include such metadata in their GEO records [<a href="#ref-4">4</a>].

Quality Control Considerations Before and After Download

Quality control is a critical step in RNA-seq analysis, and it begins with the raw data you download from SRA. Understanding the quality characteristics of your data before you invest time in alignment and quantification can save you from downstream problems.

Assessing Raw Data Quality

After converting SRA files to FASTQ, you should run a quality control tool such as FastQC to assess the quality of the raw reads. This tool generates reports on per-base quality scores, GC content, adapter contamination, and other metrics. Published RNA-seq protocols include FastQC as a standard first step after data download [<a href="#ref-2">2</a>]. The quality report helps you decide whether trimming is necessary and whether any samples should be excluded from analysis.

Identifying Adapter Contamination

Adapter contamination is a common issue in RNA-seq data, particularly when fragment sizes are shorter than the read length. The quality control report will indicate whether adapter sequences are present in your reads. If adapter contamination is detected, you will need to trim the adapters before alignment. Tools such as Trimmomatic are commonly used for this purpose [<a href="#ref-2">2</a>][<a href="#ref-3">3</a>].

Checking for Sample Mix-Ups

If you are downloading data from a public repository, you should verify that the data matches the sample descriptions in the associated publication or GEO record. This is particularly important when working with datasets generated by other laboratories. Check the number of reads, read length, and sequencing platform against the reported values. Discrepancies may indicate sample mislabeling or data processing errors.

Common Failure Patterns and Troubleshooting

Even with careful attention to the download process, you may encounter problems. Understanding common failure patterns and their solutions will help you resolve issues quickly.

Network and Connection Errors

Download failures are often caused by network issues. The prefetch command supports resumable downloads, so you can simply rerun the command if a download is interrupted. If you experience persistent connection errors, try the following:

  • Check your internet connection and firewall settings
  • Use the --max-size flag to limit the size of files downloaded in a single operation
  • Try downloading during off-peak hours when NCBI servers may be less congested
  • Consider using Aspera for large files if standard downloads are too slow

Insufficient Disk Space

SRA files and FASTQ files require substantial disk space. A single RNA-seq run can produce several gigabytes of FASTQ data, and a study with many samples can require hundreds of gigabytes. Before downloading, check the size of the dataset and ensure you have adequate storage. The prefetch command reports the size of files as it downloads them, and you can use the --max-size flag to prevent downloads that exceed your available space.

Permission Errors

If you encounter permission errors when writing files, check that you have write access to the output directories. On shared systems, you may need to use directories within your home directory or request permission from the system administrator. The SRA Toolkit configuration file may also need to be updated to point to directories where you have write access.

Paired-End File Mismatches

If the two FASTQ files for paired-end data have different numbers of reads, this indicates a problem with the conversion or the original data. Verify that you used the correct accession and that the SRA file is complete. If the problem persists, you may need to redownload the SRA file and repeat the conversion.

Configuration Issues

The SRA Toolkit uses a configuration file that controls various settings, including the location of the download cache and the behavior of certain commands. If you encounter unexpected behavior, you can reset the configuration using the vdb-config --restore command. This restores the default settings and can resolve issues caused by incorrect configuration.

Performance Optimization for Large Datasets

Downloading and converting RNA-seq data can be time-consuming, especially for large studies. Several strategies can help you optimize performance.

Parallel Downloads

The prefetch command can download multiple files simultaneously using the --max-size and --min-size flags to control the download queue. You can also run multiple prefetch commands in parallel using shell background processes or a job scheduler. However, be aware that parallel downloads may saturate your network connection and could be slower than sequential downloads in some cases.

Multithreaded Conversion

The fasterq-dump command supports multithreading with the --threads flag. Using multiple threads can significantly speed up the conversion of large SRA files to FASTQ. For example:

fasterq-dump SRR12345678 --threads 4

This uses four threads for the conversion. The optimal number of threads depends on your system's CPU resources. Published pipelines have demonstrated the use of multithreading in RNA-seq data processing [<a href="#ref-5">5</a>].

Using Cloud-Based Resources

If your local hardware is insufficient for downloading and processing large datasets, consider using cloud-based platforms. Published protocols have demonstrated the use of Google Colab for running RNA-seq analysis pipelines, including the download of raw data from NCBI GEO and SRA [<a href="#ref-2">2</a>]. Cloud platforms offer the advantage of high-bandwidth connections to NCBI servers and scalable computational resources.

Integrating SRA Downloads into Automated Pipelines

For researchers working with multiple datasets or conducting large-scale analyses, integrating SRA downloads into automated pipelines can save time and reduce errors. Several approaches are available, ranging from simple shell scripts to sophisticated workflow managers.

Shell Scripts for Batch Downloads

A simple shell script can automate the download and conversion of multiple samples. For example, you can create a text file listing all the accessions you need and use a loop to process each one:

while read accession, do
    prefetch $accession
    fasterq-dump $accession --outdir fastq --gzip
done < accessions.txt

This approach is straightforward and works well for small to medium-sized datasets.

Workflow Managers

For more complex analyses, workflow managers such as Snakemake or Nextflow provide structured ways to define and execute analysis pipelines. The nf-core project provides documentation and standards for building reproducible bioinformatics pipelines using Nextflow [<a href="#ref-6">6</a>]. These pipelines can include SRA download as an initial step, followed by quality control, alignment, and quantification. Using a workflow manager ensures that each step is executed in the correct order and that the analysis is reproducible.

Web-Based Platforms

Several web-based platforms provide automated RNA-seq analysis starting from raw SRA data. The Transcriptome-wide High-throughput RNA-seq Analysis Integrated System (THRAISE) is a web platform that performs the entire analysis pipeline, including quality control, alignment, quantification, differential expression, and functional enrichment, starting from raw NCBI SRA data [<a href="#ref-7">7</a>]. Similarly, the Confidence web application provides cross-platform differential gene expression analysis with a user-friendly interface [<a href="#ref-8">8</a>]. These platforms are useful for researchers who prefer not to work with command-line tools.

Reproducibility and Documentation Practices

Reproducibility is a fundamental principle of scientific research, and it applies to the data download process as well as to the analysis itself. Documenting your data retrieval process ensures that others can understand and reproduce your work.

Recording Software Versions

Record the version of the SRA Toolkit you used for downloading and converting data. Software updates can change behavior, and knowing the exact version helps others reproduce your results. The prefetch --version and fasterq-dump --version commands provide this information.

Documenting Accession Numbers

Maintain a complete list of all accession numbers used in your analysis, including the BioProject, SRA study, sample, and run accessions. This information should be included in your methods section when you publish your results. Published studies routinely include SRA accession numbers in their data availability statements [<a href="#ref-4">4</a>].

Saving Configuration Files

If you use non-default settings for the SRA Toolkit, document these settings in your methods. This includes the output directory structure, the use of gzip compression, and any flags used to modify the behavior of prefetch or fasterq-dump.

Using Version Control

For larger projects, consider using version control systems such as Git to track changes to your analysis scripts and configuration files. The Carpentries provides lessons on version control with Git that are relevant to scientific computing [<a href="#ref-9">9</a>]. Version control helps you track changes over time and collaborate with others more effectively.

Limitations of SRA Data and Download Tools

Understanding the limitations of SRA data and the download tools helps you interpret your results correctly and avoid common pitfalls.

Data Quality Variability

The quality of data in SRA varies widely between studies. Some datasets may have poor sequencing quality, high adapter contamination, or other issues that affect downstream analysis. You should always perform quality control on downloaded data before proceeding with analysis, regardless of the source.

Metadata Completeness

The metadata associated with SRA records is not always complete or accurate. Sample descriptions may be vague, and experimental details may be missing. When working with public datasets, you should cross-reference the SRA records with the associated publication and GEO records to obtain as much information as possible.

File Size and Storage Requirements

RNA-seq datasets are large, and downloading them requires substantial bandwidth and storage. A single RNA-seq run can produce several gigabytes of FASTQ data, and a study with many samples can require hundreds of gigabytes. Before downloading a large dataset, verify that you have adequate storage and bandwidth.

Platform-Specific Differences

The SRA Toolkit behaves slightly differently on different operating systems. Windows users may encounter path-related issues that are not present on Linux or macOS systems. The use of Windows Subsystem for Linux can help mitigate these issues [<a href="#ref-2">2</a>].

Professional Escalation Criteria

While the SRA Toolkit is designed to be reliable, there are situations where you should seek additional help. Knowing when to escalate a problem can save you time and frustration.

When to Consult NCBI Support

If you encounter persistent errors that you cannot resolve through troubleshooting, consult the NCBI SRA documentation and support resources. The NCBI website provides documentation for the SRA Toolkit and a help desk for technical issues [<a href="#ref-1">1</a>]. Include the exact error messages and the commands you used when seeking support.

When to Seek Institutional Support

If you are working in a research institution, your institution may have bioinformatics support staff who can help with data download and analysis issues. They may also have access to high-performance computing resources that can handle large datasets more efficiently.

When to Consider Alternative Data Sources

If you are unable to download data from SRA due to persistent technical issues, consider alternative sources. Some datasets are also available through the European Bioinformatics Institute's resources, which provide training and support for data retrieval [<a href="#ref-10">10</a>]. The European Nucleotide Archive (ENA) mirrors much of the data in SRA and may provide an alternative download route.

At a Glance: SRA Toolkit Commands and Their Uses

The following table summarizes the key commands covered in this tutorial and their primary uses.

CommandPrimary UseKey FlagsOutput
prefetchDownload SRA files from NCBI--output-directory, --max-size, --transportSRA files in cache or specified directory
fasterq-dumpConvert SRA files to FASTQ format--outdir, --gzip, --threads, --split-filesFASTQ files (one or two per sample)
fastq-dumpAlternative conversion tool with additional options--split-files, --gzip, --clipFASTQ files
vdb-validateVerify SRA file integrityNone requiredValidation report
vdb-configConfigure SRA Toolkit settings--restore, --setConfiguration changes

Comparison of Download Methods

The following table compares different approaches to obtaining RNA-seq data from public repositories.

MethodSpeedTechnical Skill RequiredBest For
SRA Toolkit prefetch and fasterq-dumpModerate to fastModerate command-line skillsResearchers who need local FASTQ files for custom analysis
Aspera transferFastModerate, requires Aspera clientLarge datasets where download speed is critical
Web-based platforms (THRAISE, Confidence)VariableMinimalResearchers who prefer graphical interfaces and automated pipelines
Cloud-based analysis (Google Colab)Fast with high bandwidthModerateResearchers with limited local hardware [<a href="#ref-2">2</a>]

Practical Implementation Steps

The following steps outline a practical approach to downloading RNA-seq data using the SRA Toolkit. These steps assume you have already identified the dataset you need and have the accession numbers.

Step 1: Install the SRA Toolkit

Download and install the SRA Toolkit for your operating system. Verify the installation by running prefetch --version. Configure the toolkit using vdb-config if prompted.

Step 2: Create a Project Directory Structure

Create directories for your raw SRA files and FASTQ output. A suggested structure is:

mkdir -p project/raw_sra project/fastq

Step 3: Download SRA Files

Use prefetch to download the SRA files for your accessions. For a single run:

prefetch SRR12345678 --output-directory project/raw_sra

For multiple runs, list them in a text file and use the --option-file flag.

Step 4: Validate the Downloaded Files

Run vdb-validate on each downloaded SRA file to verify integrity:

vdb-validate project/raw_sra/SRR12345678/SRR12345678.sra

Step 5: Convert SRA Files to FASTQ

Use fasterq-dump to convert the SRA files to FASTQ format:

fasterq-dump project/raw_sra/SRR12345678/SRR12345678.sra --outdir project/fastq --gzip --threads 4

Step 6: Verify the FASTQ Files

Check the number of reads in each FASTQ file and verify that paired-end files have matching read counts. Use grep -c "^@" to count reads.

Step 7: Document Your Work

Record the software version, accession numbers, and any non-default settings in your laboratory notebook or analysis documentation.

Records and Measurements to Maintain

Maintaining accurate records of your data download process is essential for reproducibility and for troubleshooting issues that may arise later in your analysis.

Download Logs

The SRA Toolkit does not automatically create detailed logs of download activity, but you can redirect command output to a log file:

prefetch SRR12345678 --output-directory project/raw_sra > download.log 2>&1

This captures both standard output and error messages, providing a record of the download process.

File Size and Read Count Records

Record the file sizes and read counts for each sample. This information is useful for verifying that downloads are complete and for planning storage requirements. The ls -lh command shows file sizes, and grep -c "^@" counts reads in FASTQ files.

Checksum Records

If you compute checksums for your downloaded files, record them in a file for future reference. This allows you to verify data integrity at any point in your analysis.

Common Failure Patterns and Their Solutions

The following table summarizes common problems encountered during SRA data download and their solutions.

ProblemLikely CauseSolution
Download fails or times outNetwork issues, server congestionRerun prefetch, try off-peak hours, use Aspera
Insufficient disk spaceDataset larger than expectedCheck dataset size before downloading, free up space, use --max-size
Permission denied errorsLack of write access to output directoryUse a directory where you have write access, check directory permissions
Paired-end files have different read countsIncomplete download or conversion errorRedownload the SRA file, rerun fasterq-dump, verify with vdb-validate
fasterq-dump produces no outputSRA file not found in expected locationSpecify the full path to the SRA file, check the cache directory
Configuration errorsIncorrect settings in vdb-configRun vdb-config --restore to reset to defaults

Safety and Ethical Considerations

Downloading and using public sequencing data carries certain responsibilities that researchers should understand.

Data Usage Policies

Public data repositories such as NCBI SRA have usage policies that govern how data can be used. Some datasets, particularly those involving human subjects, may have restrictions on data access and use. You should review the data usage policies for any dataset before downloading and using it. The NCBI website provides information about data usage policies [<a href="#ref-1">1</a>].

Attribution and Citation

When you use public datasets in your research, you should cite the original studies that generated the data and acknowledge the repositories that provide access. Published studies typically include data availability statements that specify how the data should be cited [<a href="#ref-4">4</a>].

Privacy and Confidentiality

Some sequencing datasets may contain sensitive information, particularly those derived from human subjects. Even when data is publicly available, you should handle it with care and respect the privacy of the individuals whose data is represented. De-identified data should not be re-identified or used in ways that could harm the data subjects.

Frequently Asked Questions

What is the difference between prefetch and fasterq-dump?

prefetch downloads SRA files from NCBI servers to your local machine in a compressed format. fasterq-dump converts those SRA files into FASTQ format, which is the standard input for most downstream analysis tools. You typically run prefetch first to download the data, then fasterq-dump to convert it to a usable format.

How do I know if my data is paired-end or single-end?

The SRA record for each run indicates whether the sequencing was paired-end or single-end. You can also determine this from the output of fasterq-dump. If the command produces two FASTQ files with _1 and _2 suffixes, the data is paired-end. If it produces a single file, the data is single-end.

Can I download data directly as FASTQ without using prefetch?

Yes, fasterq-dump can download and convert data in a single step if you provide the accession number directly. However, using prefetch first is often more reliable because it allows you to verify the integrity of the downloaded SRA file before conversion. The two-step approach also allows you to keep the SRA files as a backup.

How much disk space do I need for RNA-seq data?

The disk space required depends on the size of the dataset. A single RNA-seq run can produce several gigabytes of FASTQ data, and a study with many samples can require hundreds of gigabytes. Check the dataset size before downloading and ensure you have adequate storage. Using gzip compression can reduce the storage requirement.

What should I do if the download is slow?

Several strategies can improve download speed. Use the --transport aspera flag if you have the Aspera client installed. Download during off-peak hours when NCBI servers may be less congested. Use multiple threads with fasterq-dump to speed up conversion. For very large datasets, consider using cloud-based resources with high-bandwidth connections to NCBI.

How do I verify that my downloaded data is complete?

Use vdb-validate to check the integrity of SRA files. After conversion to FASTQ, count the number of reads in each file and verify that paired-end files have matching read counts. Compare the read counts and file sizes with the values reported in the SRA record.

Can I use SRA Toolkit on Windows?

Yes, the SRA Toolkit provides Windows binaries. However, many bioinformatics tools used in downstream analysis are designed for Unix-like systems. Using Windows Subsystem for Linux (WSL) allows you to run the SRA Toolkit and other bioinformatics tools in a Linux environment on Windows [<a href="#ref-2">2</a>].

What is the best way to download data for a large study with many samples?

For large studies, create a text file listing all the accession numbers and use the --option-file flag with prefetch to download them in a batch. You can also write a shell script that loops through the accessions and processes each one. For very large datasets, consider using a workflow manager such as Nextflow, which can parallelize downloads and track progress [<a href="#ref-6">6</a>].

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

[1] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information. [2] [Streamline Protocol for Bulk-RNA Sequencing: From Data Extraction to Expression Analysis.](https://pubmed.ncbi.nlm.nih.gov/41637157). Current protocols, 2026. [3] [Octopus-toolkit: a workflow to automate mining of public epigenomic and transcriptomic next-generation sequencing data.](https://pubmed.ncbi.nlm.nih.gov/29420797). Nucleic acids research, 2018. [4] [RNA-seq data exploration after trypanosome RNA-binding protein UBP1 expression is altered by CRISPR-Cas9 gene editing and overexpression](https://doi.org/10.1016/j.dib.2024.110156). Data in Brief, 2024. [5] [Development and validation of the PipeSeq program for RNA-seq data analysis in the Chlamydomonas reinhardtii as a model.](https://doi.org/10.18699/vjgb-26-34). 2026. [6] [nf-core Documentation](https://nf-co.re/docs). nf-core. [7] [THRAISE: An automated and reproducible web platform for RNA-seq analysis.](https://doi.org/10.1016/j.mocell.2025.100299). 2025. [8] [Confidence: a web app for cross-platform differential gene expression analysis, gene scoring, and enrichment analysis.](https://doi.org/10.1038/s41598-026-50527-w). 2026. [9] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries. [10] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.