How to Find and Download TCGA RNA-seq Data for Cancer Research: A Practical Guide

By Dr. Zubair Khalid, DVM, MS, PhD ·

How to Find and Download TCGA RNA-seq Data for Cancer Research: A Practical Guide

Key Takeaways

  • TCGA RNA-seq data are primarily accessed via the Genomic Data Commons (GDC) Data Portal, which offers filtering by project (cancer type), data category (e.g., Transcriptome Profiling), data type (e.g., Gene Expression Quantification), experimental strategy (RNA-Seq), and workflow type (e.g., STAR-Counts for raw counts, STAR-FPKM for normalized values).
  • Researchers must choose between raw gene counts (suitable for differential expression analysis with tools like DESeq2 or edgeR) and normalized expression values (FPKM/TPM, for cross-gene comparisons), deciding this before download to avoid reacquisition.
  • Programmatic access via the TCGAbiolinks R package is recommended for reproducible workflows, enabling scripted queries and direct integration with downstream Bioconductor analysis packages, contrasting with the GDC Data Portal's manual interface or the GDC Data Transfer Tool for large downloads.
  • Post-download quality control is critical, involving verification of sample counts against published literature (e.g., 569 samples for TCGA-UCEC), checking gene annotations against reference genomes, assessing data completeness for missing values, and confirming alignment of sample identifiers between expression and clinical data.
  • Reproducibility is paramount; document all data acquisition steps meticulously, including project code, filters applied, workflow type, data version, number of files, download date, and access method, ideally using scripted workflows with version control.

Cancer researchers frequently need to access The Cancer Genome Atlas (TCGA) RNA-seq data for studies ranging from differential expression analysis to prognostic model development. The Genomic Data Commons (GDC) Data Portal serves as the primary distribution point for TCGA data, yet many researchers find the interface overwhelming when they first encounter it. This guide provides a direct workflow for locating, filtering, and downloading TCGA RNA-seq data, with attention to the decisions that affect downstream analysis quality.

The practical outcome of this guide is a reproducible method for obtaining TCGA RNA-seq data that matches your specific cancer type, sample type, and data level requirements. You will learn how to navigate the GDC Data Portal, how to use the TCGAbiolinks R package as an alternative access route, and how to document your data acquisition steps so that your analysis remains reproducible.

Scope and Reader Context

This guide serves biology students, researchers, laboratory professionals, and life-science practitioners who need TCGA RNA-seq data for their projects. The focus is on the practical steps required to find and download data, not on the biological interpretation of any particular cancer type. The workflows described here apply across cancer types, including the projects referenced throughout the published literature, such as TCGA-UCEC for endometrial carcinoma, TCGA-OV for ovarian cancer, TCGA-STAD for gastric cancer, TCGA-KIRC for kidney renal clear cell carcinoma, TCGA-HNSC for head and neck squamous cell carcinoma, and TCGA-LUAD for lung adenocarcinoma.

The methods described in this guide have been used in published studies that integrate TCGA bulk RNA-seq data with single-cell RNA-seq data, spatial transcriptomics, and clinical information. For example, one study downloaded RNA-seq FPKM data and clinical data from the TCGA-UCEC project to analyze the tumor microenvironment in endometrial carcinoma [<a href="#ref-1">1</a>]. Another study downloaded bulk RNA-seq data from TCGA and HCCDB repositories to investigate microvascular invasion in hepatocellular carcinoma [<a href="#ref-2">2</a>]. These examples illustrate the range of research questions that begin with the same fundamental step: obtaining TCGA RNA-seq data.

Understanding TCGA RNA-seq Data Types and Levels

TCGA RNA-seq data are available in several forms, and the choice among them affects your downstream analysis. The GDC Data Portal distributes data that have been processed through standardized pipelines, and the level of processing determines what you can do with the files.

Raw Counts versus Normalized Expression

The most common distinction is between raw counts and normalized expression values. Raw counts represent the number of sequencing reads that map to each gene and are appropriate for differential expression analysis using tools such as edgeR or DESeq2. Normalized values, such as FPKM (fragments per kilobase of transcript per million mapped fragments) or TPM (transcripts per million), adjust for gene length and sequencing depth and are appropriate for comparing expression levels across genes or samples.

Published studies have used both forms. One study of endometrial carcinoma downloaded RNA-Seq FPKM data from the TCGA-UCEC project [<a href="#ref-1">1</a>]. Another study of gastric cancer downloaded high-throughput mRNA expression data from the TCGA database and used the edgeR software package for differential analysis [<a href="#ref-3">3</a>]. The choice between raw counts and normalized values depends on your analysis plan, and you should decide before downloading to avoid reacquiring data.

Data Portals and Repositories

The GDC Data Portal is the primary distribution point for TCGA data, but it is not the only source. The Gene Expression Omnibus (GEO) at NCBI hosts many cancer RNA-seq datasets, including those generated by independent studies [<a href="#ref-4">4</a>]. The International Cancer Genome Consortium (ICGC) also provides cancer genomics data, and some studies have used ICGC cohorts for external validation of models built on TCGA data [<a href="#ref-5">5</a>].

The NCBI provides a range of data resources and search systems that support cancer research, including sequence databases and analysis services [<a href="#ref-4">4</a>]. For researchers who need training on how to use these resources effectively, the EMBL-EBI Training program offers learning pathways for bioinformatics data resources and practical analysis education [<a href="#ref-6">6</a>].

The GDC Data Portal Interface

The GDC Data Portal is the official distribution point for TCGA data. The interface allows you to filter data by project, data category, data type, experimental strategy, and other attributes. Understanding the interface structure is the first step toward efficient data retrieval.

Projects and Cohorts

Each TCGA project is identified by a code that indicates the cancer type. For example, TCGA-UCEC corresponds to uterine corpus endometrial carcinoma, TCGA-OV corresponds to ovarian cancer, TCGA-STAD corresponds to stomach adenocarcinoma, TCGA-KIRC corresponds to kidney renal clear cell carcinoma, TCGA-HNSC corresponds to head and neck squamous cell carcinoma, and TCGA-LUAD corresponds to lung adenocarcinoma.

When you enter the GDC Data Portal, you can browse projects by cancer type or search for a specific project code. Selecting a project narrows the available files to those generated for that cancer cohort.

Data Categories and Experimental Strategies

Within a project, you can filter by data category. For RNA-seq analysis, the relevant categories include transcriptome profiling and gene expression quantification. The experimental strategy should be set to RNA-seq, and you can further filter by the type of RNA that was sequenced, such as total RNA or mRNA.

The GDC Data Portal also allows you to filter by workflow type, which determines the processing pipeline applied to the raw data. Different workflows produce different output formats, including STAR-Counts for raw counts and STAR-FPKM for normalized expression values.

Sample Types and Cases

TCGA data include multiple sample types for each case, including primary tumors, normal tissue adjacent to the tumor, and metastatic tumors. Your analysis plan determines which sample types you need. For differential expression analysis between tumor and normal tissue, you need both primary tumor and solid tissue normal samples. For studies that focus exclusively on tumor biology, you may only need primary tumor samples.

The GDC Data Portal allows you to filter by sample type, and you can also filter by the number of cases to ensure that you are downloading data from a sufficient number of patients for your analysis.

Step-by-Step Workflow for the GDC Data Portal

The following workflow describes the process of finding and downloading TCGA RNA-seq data through the GDC Data Portal. The steps are presented in the order you would perform them.

Step 1: Access the GDC Data Portal

Navigate to the GDC Data Portal website. The portal is the official distribution point for TCGA data and provides access to the full range of data generated by the project.

Step 2: Select Your Project

Use the Projects page to find your cancer type of interest. You can browse the list of projects or use the search function to find a specific project code. Click on the project to view its summary page, which shows the number of cases, files, and available data categories.

Step 3: Apply Filters for RNA-seq Data

From the project page, navigate to the Files page. Apply the following filters:

  • Data Category: Transcriptome Profiling
  • Data Type: Gene Expression Quantification
  • Experimental Strategy: RNA-Seq
  • Workflow Type: Select the workflow that produces the output format you need, such as STAR-Counts for raw counts or STAR-FPKM for normalized expression

These filters narrow the file list to RNA-seq gene expression data for your selected project.

Step 4: Filter by Sample Type

Use the Cases filter to select the sample types you need. For tumor versus normal comparisons, select both Primary Tumor and Solid Tissue Normal. For tumor-only studies, select Primary Tumor.

Step 5: Review the File Manifest

After applying your filters, review the file list to confirm that the files match your requirements. The file list shows the file name, data category, data type, experimental strategy, and associated case. You can sort the list by any column to organize the files.

Step 6: Download the Files

Select the files you want to download and use the Download button. You have two options:

  • Download the files directly to your computer
  • Download a manifest file that you can use with the GDC Data Transfer Tool for large downloads

For small numbers of files, direct download is sufficient. For large cohorts with hundreds of samples, the GDC Data Transfer Tool is more reliable because it supports resumable downloads and parallel transfer.

Step 7: Document Your Download

Record the date of download, the filters you applied, the number of files downloaded, and the workflow type. This documentation is essential for reproducibility, as it allows other researchers to understand exactly which data you obtained.

Using TCGAbiolinks in R

TCGAbiolinks is an R package that provides programmatic access to TCGA data through the GDC Application Programming Interface (API). The package is distributed through Bioconductor, which provides official documentation for package installation and reproducible genomic analysis workflows [<a href="#ref-7">7</a>].

Installing TCGAbiolinks

TCGAbiolinks is installed through Bioconductor. The Bioconductor project provides official documentation for package installation and workflow development [<a href="#ref-7">7</a>]. After installation, you load the package in R and use its functions to query and download TCGA data.

Querying the GDC API

The TCGAbiolinks workflow begins with a query function that specifies the project, data category, data type, and other parameters. The query returns a list of available files that match your criteria. You can then review the query results before downloading.

Downloading Data with TCGAbiolinks

After reviewing the query results, you use the download function to retrieve the files. TCGAbiolinks downloads the files to a local directory and prepares them for analysis. The package also provides functions for preparing the data for downstream analysis, including the creation of a summarized experiment object.

Preparing Data for Analysis

TCGAbiolinks includes functions for data preparation that convert the downloaded files into formats suitable for analysis. These functions handle the mapping of file names to sample identifiers and the assembly of expression matrices. The prepared data can then be used with other Bioconductor packages for differential expression analysis, survival analysis, and other downstream applications.

Comparison of Access Methods

The choice between the GDC Data Portal and TCGAbiolinks depends on your workflow and technical preferences. The following table summarizes the key differences.

Access MethodBest ForKey AdvantageKey Limitation
GDC Data Portal Web InterfaceOne-time downloads, visual explorationNo programming required, visual filtersManual process, difficult to reproduce exactly
GDC Data Transfer ToolLarge cohort downloadsResumable downloads, parallel transferRequires command line familiarity
TCGAbiolinks in RReproducible workflows, integration with analysisScripted queries, direct integration with BioconductorRequires R programming knowledge

The GDC Data Portal web interface is appropriate for researchers who need to download data once and do not plan to repeat the process. The GDC Data Transfer Tool is appropriate for large downloads where reliability is important. TCGAbiolinks is appropriate for researchers who want to integrate data download with their analysis workflow and who value reproducibility through scripted queries.

Quality Control Considerations

Quality control is an essential step after downloading TCGA RNA-seq data. The data have been processed through standardized pipelines, but you should verify that the data meet your expectations before proceeding with analysis.

Verifying Sample Counts

After downloading, verify that the number of samples matches your expectation based on the filters you applied. Published studies report specific sample counts for their cohorts. For example, one study of endometrial carcinoma obtained 569 RNA-seq samples from the TCGA-UCEC project [<a href="#ref-1">1</a>]. Another study of gastric cancer obtained RNA-seq expression data from 375 cancer cases and 32 adjacent tissue samples [<a href="#ref-3">3</a>]. These numbers provide reference points for the expected cohort sizes.

Checking Gene Annotations

Verify that the gene annotations in your downloaded data match the reference genome and annotation version you plan to use. Mismatches between the data annotation and your analysis tools can produce errors in gene identification and quantification.

Assessing Data Completeness

Check for missing values in the expression matrix. Some genes may have zero counts in some samples, and some samples may have failed quality control in the original processing. Understanding the pattern of missing data helps you decide whether to filter genes or samples before analysis.

Confirming Clinical Data Alignment

If your analysis uses clinical data, verify that the sample identifiers in the expression data match the sample identifiers in the clinical data. Mismatched identifiers are a common source of errors in downstream analysis.

Common Failure Patterns and Troubleshooting

Researchers encounter several common problems when downloading TCGA RNA-seq data. Understanding these failure patterns helps you avoid them or resolve them quickly.

Filter Mismatches

A common problem is applying filters that are too restrictive or too permissive. For example, selecting the wrong workflow type can return files that are not compatible with your analysis tools. Review the filter settings carefully before downloading.

Incomplete Downloads

Large downloads can fail due to network interruptions or server timeouts. The GDC Data Transfer Tool provides resumable downloads that reduce the impact of these failures. For direct downloads, verify that all files were downloaded completely by comparing the file count and file sizes with the manifest.

Identifier Confusion

TCGA sample identifiers follow a specific format that encodes the project, participant, sample type, and aliquot. Confusion between sample identifiers and file identifiers can lead to errors in data assembly. Use the GDC metadata to map file names to sample identifiers correctly.

Version Inconsistencies

TCGA data are periodically updated, and different versions of the data may produce different results. Record the data version and the date of download to ensure that your analysis can be reproduced or compared with other studies.

Reproducibility and Documentation

Reproducibility is a central concern in bioinformatics analysis. The steps you take to document your data acquisition process determine whether other researchers can reproduce your work.

Recording Data Acquisition Steps

Record the following information for every TCGA RNA-seq data download:

  • Project code and cancer type
  • Data category, data type, and experimental strategy
  • Workflow type and data version
  • Sample types included
  • Number of cases and files
  • Date of download
  • Access method used

This information allows other researchers to understand exactly which data you obtained and to repeat the process if needed.

Using Scripted Workflows

Scripted workflows using TCGAbiolinks provide a higher level of reproducibility than manual downloads through the web interface. The script records the query parameters and download steps, making the process transparent and repeatable.

Version Control for Analysis Code

Store your data acquisition scripts and analysis code in a version control system. The Carpentries provides lessons on foundational computing and data skills, including version control with Git, that support reproducible research practices [<a href="#ref-8">8</a>].

Integration with Downstream Analysis

TCGA RNA-seq data are frequently used as the bulk RNA-seq component in studies that integrate multiple data types. Understanding how your downloaded data will be used downstream helps you make appropriate decisions during data acquisition.

Differential Expression Analysis

Differential expression analysis compares gene expression between conditions, such as tumor versus normal tissue. Tools such as edgeR operate on raw counts and require a count matrix as input. If you plan to perform differential expression analysis, download raw counts instead of normalized values.

One study of gastric cancer used the edgeR package to perform differential analysis on TCGA data, applying thresholds of |logFC| greater than 1 and p less than 0.05 to identify significantly expressed genes [<a href="#ref-3">3</a>]. The study identified 4320 differential genes, of which 2718 were highly expressed and 1602 were low expressed [<a href="#ref-3">3</a>].

Prognostic Model Development

Many published studies use TCGA RNA-seq data to develop prognostic models. These studies typically combine expression data with clinical data to identify genes associated with survival outcomes.

A study of triple negative breast cancer downloaded RNA-seq expression data from TCGA for 116 breast cancer samples lacking ER, PR, and HER2 expression and 113 normal tissue samples [<a href="#ref-9">9</a>]. The study used weighted gene co-expression network analysis and differential gene expression analysis to screen for differentially co-expressed genes, then applied LASSO feature selection to identify key genes [<a href="#ref-9">9</a>].

A study of clear cell renal cell carcinoma downloaded gene expression and clinical data from TCGA and GEO databases to develop a stemness-related long non-coding RNA signature [<a href="#ref-10">10</a>]. The study used weighted correlation network analysis to identify stemness-related genes and multiple machine learning algorithms to construct a prognostic signature [<a href="#ref-10">10</a>].

Integration with Single-Cell Data

TCGA bulk RNA-seq data are often integrated with single-cell RNA-seq data to connect bulk-level observations with cell-level heterogeneity. The single-cell data are typically downloaded from GEO, while the bulk data come from TCGA.

A study of intestinal-type gastric cancer downloaded single-cell RNA-seq data from GEO and used TCGA samples for immune cell infiltration analysis [<a href="#ref-11">11</a>]. The study constructed a Cox proportional risk regression model to identify prognostic genes [<a href="#ref-11">11</a>].

A study of glioblastoma integrated single-cell RNA-seq data with bulk RNA-seq data from TCGA, the Chinese Glioma Genome Atlas, and GEO to build a prognostic model [<a href="#ref-12">12</a>]. The study used the Seurat package for single-cell data processing and applied univariate Cox and LASSO analyses to screen prognostic genes [<a href="#ref-12">12</a>].

Deconvolution Analysis

Deconvolution methods estimate the cellular composition of bulk RNA-seq samples. A study evaluating deconvolution methods used 18 real bulk RNA-expression cohorts totaling 5,891 samples across nine cancer types [<a href="#ref-13">13</a>]. The study found that ReCIDE and BayesPrism were robust deconvolution methods and identified matrix cancer-associated fibroblasts as a prognostic marker with consistent effects across multiple cancers [<a href="#ref-13">13</a>].

Training and Learning Resources

Researchers who are new to TCGA data analysis can benefit from structured training resources. Several organizations provide free or open-access training materials.

Bioconductor Training

The Bioconductor project provides official documentation for packages, workflows, installation, and reproducible genomic analysis [<a href="#ref-7">7</a>]. The documentation includes vignettes and workflow articles that demonstrate how to use packages such as TCGAbiolinks for TCGA data analysis.

EMBL-EBI Training

The EMBL-EBI Training program offers learning pathways for bioinformatics data resources and practical analysis education [<a href="#ref-6">6</a>]. These resources are useful for researchers who need to build foundational skills in data analysis.

Galaxy Training Network

The Galaxy Training Network provides accessible workflow training and analysis tutorials [<a href="#ref-14">14</a>]. The tutorials cover a range of topics, including RNA-seq analysis, and emphasize reproducibility through documented workflows.

nf-core Documentation

The nf-core community provides documentation for community pipeline standards, usage, configuration, and reproducible workflow context [<a href="#ref-15">15</a>]. The pipelines are designed to be reproducible and portable across computing environments.

The Carpentries

The Carpentries provides lessons on foundational computing, data, shell, Git, and programming training [<a href="#ref-8">8</a>]. These lessons are valuable for researchers who need to build the technical skills required for bioinformatics analysis.

Published Examples of TCGA Data Usage

The published literature provides numerous examples of how TCGA RNA-seq data are used in cancer research. These examples illustrate the range of analysis approaches and the decisions researchers make during data acquisition.

Endometrial Carcinoma Studies

A study of endometrial carcinoma downloaded RNA-Seq FPKM data and clinical data from the TCGA-UCEC project [<a href="#ref-1">1</a>]. The study also downloaded single-cell RNA-seq and spatial transcriptome data from GEO and analyzed the data using R software [<a href="#ref-1">1</a>]. The analysis identified cell populations in the tumor microenvironment and examined cell-cell communication [<a href="#ref-1">1</a>].

Another study of endometrial carcinoma downloaded RNA-seq data from TCGA and identified IFN-gamma-related long non-coding RNAs using co-expression analysis and weighted co-expression network analysis [<a href="#ref-16">16</a>]. The study constructed a prognostic signature using univariate Cox regression and LASSO regression [<a href="#ref-16">16</a>].

Hepatocellular Carcinoma Studies

Multiple studies have used TCGA data to investigate hepatocellular carcinoma. One study downloaded bulk RNA-seq data from TCGA and HCCDB repositories, single-cell RNA-seq data from GEO, and spatial transcriptomics data from CNCB [<a href="#ref-2">2</a>]. The study used the Scissor algorithm to delineate prognosis-related cell subpopulations and identified a microvascular invasion-related malignant cell subtype [<a href="#ref-2">2</a>].

Another study downloaded RNA-seq information for hepatocellular carcinoma patients from TCGA and ICGC databases [<a href="#ref-5">5</a>]. The study used single-cell RNA-seq data from GSE166635 and identified T-cell exhaustion-related genes through gene set variance analysis and weighted gene correlation network analysis [<a href="#ref-5">5</a>].

A third study downloaded RNA-seq data from TCGA, GEO, and ICGC databases and acquired single-cell data from GSE149614 [<a href="#ref-17">17</a>]. The study constructed a CD8+ T-cell exhaustion signature using differential gene analysis, univariate Cox regression, LASSO regression, and multivariate Cox regression [<a href="#ref-17">17</a>].

Gastric Cancer Studies

A study of gastric cancer downloaded transcriptome profiling data from TCGA-STAD and GSE84437 and single-cell RNA-seq data from GSE167297 [<a href="#ref-18">18</a>]. The study analyzed COL5A2 expression and its relationship with clinicopathological factors, performed survival analysis and Cox regression analysis, and explored correlations with immune cell infiltration [<a href="#ref-18">18</a>].

Another study of gastric cancer downloaded high-throughput mRNA expression data from TCGA and used the edgeR package for differential analysis [<a href="#ref-3">3</a>]. The study performed enrichment analysis and constructed a protein interaction network to identify significant genes [<a href="#ref-3">3</a>].

Colorectal Cancer Studies

A study of colorectal cancer analyzed oxidative stress-related pathway activities using both single-cell and bulk RNA-seq data [<a href="#ref-19">19</a>]. The study used the TCGA-CRC dataset and identified two oxidative stress-specific clusters, then established a 12-gene risk signature using the LASSO Cox method [<a href="#ref-19">19</a>].

Another study of colorectal cancer used single-cell RNA-seq analysis to identify macrophages and classify them into M0-like, M1-like, and M2-like states [<a href="#ref-20">20</a>]. The study identified M2-associated prognostic genes and validated their robustness in an independent dataset [<a href="#ref-20">20</a>].

Other Cancer Types

A study of lung adenocarcinoma downloaded RNA-seq data from the TCGA LUAD cohort from the Genomic Data Commons Portal [<a href="#ref-21">21</a>]. The study used the GSE13213 dataset from GEO for external validation and identified lysosome-related genes associated with prognosis [<a href="#ref-21">21</a>].

A study of colon cancer downloaded RNA-seq data from the TCGA dataset and focused on DNA methylation genes in the tumor immune microenvironment [<a href="#ref-22">22</a>]. The study used the GSVA method to score DNA methylation genes and performed enrichment analysis [<a href="#ref-22">22</a>].

A study of head and neck cancer performed whole-transcriptome spatial transcriptomics on tissue sections and developed a 28-gene tumor budding signature [<a href="#ref-23">23</a>]. The signature was validated using bulk RNA-seq from TCGA-HNSC, single-cell RNA-seq, and independent spatial transcriptomics datasets [<a href="#ref-23">23</a>].

Data Management and Storage

TCGA RNA-seq datasets can be large, particularly when you download raw counts for hundreds of samples. Proper data management practices help you avoid storage problems and ensure that your data remain accessible for analysis.

Storage Requirements

Estimate the storage requirements before downloading. Raw count matrices for a typical TCGA cohort occupy several gigabytes. Normalized expression matrices are smaller but still require careful storage planning.

Directory Structure

Organize your downloaded data in a clear directory structure. A common approach is to create separate directories for raw data, processed data, and analysis results. This structure makes it easier to track the provenance of your data and to share your workflow with collaborators.

Backup and Archiving

Maintain backups of your downloaded data, particularly if the data are difficult to replace. The GDC Data Portal allows you to re-download data, but the process takes time and may be subject to changes in data versions.

Limitations of TCGA RNA-seq Data

TCGA RNA-seq data have limitations that affect their use in research. Understanding these limitations helps you interpret your results appropriately.

Sample Composition

TCGA samples are predominantly from primary tumors. Metastatic samples and samples from rare cancer subtypes are underrepresented. If your research question involves metastatic disease or a rare subtype, TCGA data may not provide sufficient sample numbers.

Batch Effects

TCGA data were generated over many years across multiple sequencing centers. Batch effects can introduce technical variation that confounds biological signals. Your analysis should account for potential batch effects, particularly when comparing samples processed at different times or locations.

Normal Tissue Availability

Not all TCGA projects include normal tissue samples. The availability of matched normal samples varies by project, and some projects have very few normal samples. This limitation affects your ability to perform tumor versus normal comparisons.

Clinical Data Completeness

Clinical data associated with TCGA samples vary in completeness. Some samples have comprehensive clinical annotations, while others have limited data. Review the clinical data availability for your project before planning analyses that depend on clinical variables.

Professional Escalation Criteria

Some situations require consultation with a bioinformatics specialist or a more experienced colleague. Recognizing these situations helps you avoid errors that could compromise your analysis.

When to Seek Help with Data Access

Consult a specialist if you encounter any of the following situations:

  • The GDC Data Portal does not return the expected files for your project
  • The downloaded files do not match the manifest or the expected file count
  • You are uncertain about which workflow type or data level is appropriate for your analysis
  • You need to download data for a project that is not listed in the GDC Data Portal

When to Seek Help with Data Analysis

Consult a specialist if you encounter any of the following situations:

  • The expression matrix contains unexpected patterns of missing values
  • The sample identifiers in the expression data do not match the clinical data
  • The results of your quality control checks indicate problems with the data
  • You are uncertain about how to account for batch effects in your analysis

When to Seek Help with Interpretation

Consult a specialist if you encounter any of the following situations:

  • Your analysis produces results that contradict established findings in the literature
  • You are uncertain about the biological interpretation of your results
  • You need to integrate TCGA data with other data types and are uncertain about the appropriate methods

Records and Documentation Practices

Maintaining accurate records of your data acquisition process is essential for reproducible research. The following practices support good record keeping.

Download Logs

Create a download log that records the date, project, filters, file count, and access method for each download. This log serves as a permanent record of your data acquisition process.

Analysis Scripts

Store all analysis scripts in a version control system. The Carpentries provides lessons on version control with Git that support reproducible research practices [<a href="#ref-8">8</a>].

Metadata Files

Retain the metadata files that accompany your downloaded data. These files contain information about sample annotations, clinical data, and processing details that are essential for analysis.

README Files

Create a README file for each project that describes the data acquisition process, the analysis workflow, and any decisions that affected the analysis. This documentation helps you and your collaborators understand the context of your results.

Safety and Ethical Considerations

TCGA data include clinical information from cancer patients. Although the data are de-identified, you should handle them with appropriate care and follow the data use agreements that apply to TCGA data.

Data Use Agreements

TCGA data are available for research use under the data use agreements established by the National Institutes of Health. Review the terms of these agreements before downloading and using the data.

Responsible Data Handling

Store TCGA data securely and limit access to authorized members of your research team. Do not attempt to re-identify patients from the data, and do not share the data in ways that violate the data use agreements.

Publication Practices

When you publish results based on TCGA data, cite the data source appropriately and describe your data acquisition process in the methods section of your paper. This practice supports the reproducibility of your work and gives credit to the researchers who generated the data.

At a Glance

The following table summarizes the key decisions you need to make when downloading TCGA RNA-seq data.

Decision PointOptionsConsideration
Data LevelRaw counts or normalized expression (FPKM, TPM)Raw counts for differential expression, normalized values for cross-gene comparisons
Access MethodGDC Data Portal web interface, GDC Data Transfer Tool, TCGAbiolinksWeb interface for one-time downloads, transfer tool for large downloads, TCGAbiolinks for reproducible workflows
Sample TypesPrimary tumor, solid tissue normal, metastaticTumor and normal for differential expression, tumor only for tumor biology studies
Workflow TypeSTAR-Counts, STAR-FPKM, other pipelinesDetermines the output format and compatibility with analysis tools
DocumentationDownload log, metadata files, analysis scriptsEssential for reproducibility and publication

Practical Implementation Steps

The following steps provide a practical implementation plan for downloading TCGA RNA-seq data.

Step 1: Define Your Analysis Requirements

Before downloading data, define your analysis requirements. Determine the cancer type, the sample types, the data level, and the number of samples you need. This planning step prevents unnecessary downloads and ensures that you obtain the appropriate data.

Step 2: Choose Your Access Method

Select the access method that matches your technical skills and workflow. If you are comfortable with R programming, TCGAbiolinks provides a reproducible approach. If you prefer a visual interface, the GDC Data Portal web interface is appropriate.

Step 3: Download the Data

Execute your chosen access method and download the data. Verify that the downloaded files match your expectations in terms of file count and file sizes.

Step 4: Perform Quality Control

Run quality control checks on the downloaded data. Verify sample counts, check gene annotations, assess data completeness, and confirm clinical data alignment.

Step 5: Document Your Process

Record the details of your data acquisition process, including the date, filters, file counts, and access method. Store this documentation with your analysis scripts.

Step 6: Prepare Data for Analysis

Convert the downloaded data into the format required by your analysis tools. For differential expression analysis, create a count matrix. For other analyses, prepare the appropriate expression matrix.

Observations and Measurements

The published literature provides reference points for the scale of TCGA RNA-seq datasets. These observations help you calibrate your expectations when downloading data.

Cohort Sizes

TCGA cohorts vary in size by cancer type. A study of endometrial carcinoma obtained 569 RNA-seq samples from the TCGA-UCEC project [<a href="#ref-1">1</a>]. A study of gastric cancer obtained RNA-seq expression data from 375 cancer cases and 32 adjacent tissue samples [<a href="#ref-3">3</a>]. A study of triple negative breast cancer downloaded data from 116 cancer samples and 113 normal tissue samples [<a href="#ref-9">9</a>].

Gene Detection

The number of genes detected in TCGA RNA-seq data depends on the sequencing depth and the analysis pipeline. A study of endometrial carcinoma detected 33,408 genes across 33,162 cells in single-cell RNA-seq data [<a href="#ref-1">1</a>]. The number of genes detected in bulk RNA-seq data is typically lower but still in the range of 20,000 to 30,000.

Differential Gene Counts

Differential expression analysis of TCGA data typically identifies thousands of differentially expressed genes. A study of gastric cancer identified 4,320 differential genes, of which 2,718 were highly expressed and 1,602 were low expressed [<a href="#ref-3">3</a>].

Common Failure Patterns in TCGA Data Analysis

Beyond the data acquisition failures described earlier, researchers encounter several common failure patterns in TCGA data analysis.

Overlooking Batch Effects

TCGA data were generated across multiple sequencing centers over many years. Failing to account for batch effects can produce spurious results. Include batch information in your analysis model when appropriate.

Ignoring Data Version Differences

TCGA data are periodically updated, and different versions may produce different results. Always record the data version and date of download to ensure that your analysis can be reproduced.

Misaligning Sample Identifiers

Sample identifiers in the expression data must match the sample identifiers in the clinical data. Misalignment produces errors in survival analysis and other clinical correlations.

Using Inappropriate Normalization

The choice of normalization method affects the results of your analysis. Raw counts are appropriate for differential expression analysis with tools such as edgeR. Normalized values are appropriate for comparing expression levels across genes.

Frequently Asked Questions

What is the difference between raw counts and FPKM values in TCGA RNA-seq data?

Raw counts represent the number of sequencing reads that map to each gene and are appropriate for differential expression analysis using tools such as edgeR or DESeq2. FPKM values are normalized for gene length and sequencing depth and are appropriate for comparing expression levels across genes or samples. The choice between them depends on your analysis plan. For differential expression analysis, download raw counts. For cross-gene comparisons, download FPKM or TPM values.

How do I choose between the GDC Data Portal and TCGAbiolinks for downloading TCGA data?

The GDC Data Portal web interface is appropriate for one-time downloads and for researchers who prefer a visual interface. TCGAbiolinks is appropriate for researchers who want to integrate data download with their analysis workflow and who value reproducibility through scripted queries. The GDC Data Transfer Tool is appropriate for large downloads where resumable downloads and parallel transfer are important.

Which TCGA projects are available for RNA-seq data download?

TCGA includes projects for many cancer types, each identified by a project code. Examples include TCGA-UCEC for uterine corpus endometrial carcinoma, TCGA-OV for ovarian cancer, TCGA-STAD for stomach adenocarcinoma, TCGA-KIRC for kidney renal clear cell carcinoma, TCGA-HNSC for head and neck squamous cell carcinoma, and TCGA-LUAD for lung adenocarcinoma. The GDC Data Portal provides a complete list of available projects.

How do I filter TCGA data to include only primary tumor samples?

In the GDC Data Portal, use the Cases filter to select the sample type. For primary tumor samples, select Primary Tumor in the sample type filter. You can also select Solid Tissue Normal for normal samples adjacent to the tumor. The filter narrows the file list to the selected sample types.

What quality control checks should I perform after downloading TCGA RNA-seq data?

Verify that the number of samples matches your expectation based on the filters you applied. Check that the gene annotations match the reference genome and annotation version you plan to use. Assess the pattern of missing values in the expression matrix. Confirm that the sample identifiers in the expression data match the sample identifiers in the clinical data.

How do I document my TCGA data download for reproducibility?

Record the project code, data category, data type, experimental strategy, workflow type, sample types, number of cases and files, date of download, and access method. Store this information in a download log or README file. If you use TCGAbiolinks, the script itself serves as documentation of the query parameters and download steps.

Can I use TCGA RNA-seq data for differential expression analysis between tumor and normal samples?

Yes, TCGA RNA-seq data are commonly used for differential expression analysis between tumor and normal samples. You need to download data for both primary tumor and solid tissue normal samples. Use raw counts for differential expression analysis with tools such as edgeR. Note that not all TCGA projects include normal tissue samples, and the availability of matched normal samples varies by project.

How do I integrate TCGA bulk RNA-seq data with single-cell RNA-seq data?

TCGA bulk RNA-seq data are frequently integrated with single-cell RNA-seq data to connect bulk-level observations with cell-level heterogeneity. The single-cell data are typically downloaded from GEO, while the bulk data come from TCGA. Published studies have used this approach for endometrial carcinoma [<a href="#ref-1">1</a>], hepatocellular carcinoma [<a href="#ref-2">2</a>], gastric cancer [<a href="#ref-11">11</a>], and glioblastoma [<a href="#ref-12">12</a>]. The integration typically involves analyzing the single-cell data to identify cell populations and then using the bulk data to validate or extend the findings.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

[1] [Integrating single-cell RNA-seq and spatial transcriptomics reveals MDK-NCL dependent immunosuppressive environment in endometrial carcinoma.](https://pubmed.ncbi.nlm.nih.gov/37081869). Frontiers in immunology, 2023. [2] [Multi-transcriptomics analysis of microvascular invasion-related malignant cells and development of a machine learning-based prognostic model in hepatocellular carcinoma.](https://pubmed.ncbi.nlm.nih.gov/39176099). Frontiers in immunology, 2024. [3] [Analysis and Prediction of Significant Genes in Gastric Cancer Based on TCGA Database.](https://doi.org/10.3233/SHTI230871). Studies in Health Technology and Informatics, 2023. [4] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information. [5] [T-cell exhaustion signatures characterize the immune landscape and predict HCC prognosis via integrating single-cell RNA-seq and bulk RNA-sequencing.](https://pubmed.ncbi.nlm.nih.gov/37006257). Frontiers in immunology, 2023. [6] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute. [7] [Bioconductor](https://bioconductor.org/). Bioconductor Project. [8] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries. [9] [Identification of Key Prognostic Genes of Triple Negative Breast Cancer by LASSO-Based Machine Learning and Bioinformatics Analysis.](https://pubmed.ncbi.nlm.nih.gov/35627287). Genes, 2022. [10] [A novel stemness-related lncRNA signature predicts prognosis, immune infiltration and drug sensitivity of clear cell renal cell carcinoma.](https://pubmed.ncbi.nlm.nih.gov/40016772). Journal of translational medicine, 2025. [11] [Integrated analysis of single-cell RNA-seq and bulk RNA-seq to unravel the molecular mechanisms underlying the immune microenvironment in the development of intestinal-type gastric cancer.](https://pubmed.ncbi.nlm.nih.gov/37591405). Biochimica et biophysica acta. Molecular basis of disease, 2024. [12] [Joint analysis of single-cell RNA sequencing and bulk transcriptome reveals the heterogeneity of the urea cycle of astrocytes in glioblastoma.](https://pubmed.ncbi.nlm.nih.gov/39938577). Neurobiology of disease, 2025. [13] [Evaluating deconvolution methods using real bulk RNA-expression data for robust prognostic insights across cancer types.](https://doi.org/10.1186/s13059-026-03942-1). 2026. [14] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project. [15] [nf-core Documentation](https://nf-co.re/docs). nf-core. [16] [The IFN-γ-related long non-coding RNA signature predicts prognosis and indicates immune microenvironment infiltration in uterine corpus endometrial carcinoma](https://doi.org/10.3389/fonc.2022.955979). Frontiers in Oncology, 2022. [17] [Identification of the CD8+ T-cell exhaustion signature of hepatocellular carcinoma for the prediction of prognosis and immune microenvironment by integrated analysis of bulk- and single-cell RNA sequencing data](https://doi.org/10.21037/tcr-24-650). Translational Cancer Research, 2023. [18] [COL5A2 is a prognostic-related biomarker and correlated with immune infiltrates in gastric cancer based on transcriptomics and single-cell RNA sequencing](https://doi.org/10.1186/s12920-023-01659-9). BMC Medical Genomics, 2023. [19] [Deciphering oxidative stress-related heterogeneity and developing a prognostic signature for colorectal cancer.](https://doi.org/10.1038/s41598-026-46560-4). 2026. [20] [M2 macrophage associated genes shape prognosis and tumor progression in human colorectal cancer.](https://doi.org/10.1016/j.isci.2026.116230). 2026. [21] [Identification of lysosomal genes associated with prognosis in lung adenocarcinoma](https://doi.org/10.21037/tlcr-23-14). Translational Lung Cancer Research, 2023. [22] [LARS2 DNA methylation predicts the prognosis in colon cancer](https://doi.org/10.1038/s41598-025-10669-9). Scientific Reports, 2025. [23] [Spatial transcriptomics reveals a molecular tumor budding signature in head and neck cancer.](https://doi.org/10.1186/s13073-026-01612-2). 2026.

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.