# RNA-seq Count Data: From Raw Counts to Differential Expression - A Complete Workflow with Galaxy


## Key Takeaways

- The Galaxy platform facilitates RNA-seq differential expression analysis for researchers lacking programming skills by integrating multiple bioinformatics tools via a web-based graphical interface, enabling a complete workflow from raw FASTQ files to statistical interpretation.
- A reference-based workflow necessitates a well-annotated genome (FASTA) and gene annotation (GTF/GFF); quality control of raw reads using tools like FastQC and adapter trimming with Trimmomatic are critical initial steps to ensure data integrity.
- Read alignment, typically performed with splice-aware aligners such as HISAT2, maps sequencing reads to the reference genome, with key parameters including mismatch allowance and multi-mapping read handling influencing downstream quantification accuracy.
- Gene expression quantification, commonly achieved with FeatureCounts, generates a count matrix by assigning aligned reads to genomic features, which then serves as input for differential expression analysis using statistical methods like DESeq2.
- DESeq2 models count data with a negative binomial distribution to identify differentially expressed genes, accounting for biological variability and using adjusted p-values to control the false discovery rate, with results often validated by independent methods like RT-qPCR.
- Reproducibility is paramount, with Galaxy automatically recording analysis history; essential documentation includes tool versions, exact parameters, reference genome/annotation versions, and key measurements such as mapping rates and read counts per sample.

---

RNA sequencing has become the standard method for transcriptome analysis, with gene expression quantification serving as the core application in molecular biology research. For researchers without programming skills, the computational complexity of RNA-seq analysis presents a significant barrier. The Galaxy platform addresses this challenge by providing a web-based environment that integrates multiple bioinformatics tools and databases, enabling researchers without extensive bioinformatics expertise to effectively analyze RNA-seq data. This article walks through a complete Galaxy workflow from raw FASTQ files to differential expression results, including tool choices and parameter settings that affect downstream interpretation.

## Scope and Reader Context

This workflow is designed for biology students, researchers, laboratory professionals, and life-science practitioners who need to perform differential expression analysis without writing code. The reference-based approach described here assumes a well-annotated genome is available for the organism under study. The Galaxy public web server provides the computational infrastructure, removing the need for local high-performance computing resources. The workflow covers quality control, read alignment, quantification, and statistical analysis of differential expression, with attention to the decisions that influence biological conclusions.

## The RNA-seq Analysis Problem

RNA-seq data analysis presents challenges that include complex computational pipelines, the need for bioinformatics skills, and significant computational requirements. These obstacles historically excluded many laboratory scientists from performing their own transcriptome analyses. The Galaxy platform solves this problem by offering an easy-to-use, web-based solution that integrates multiple bioinformatics tools and databases. This enables researchers without extensive bioinformatics expertise to effectively analyze RNA-seq data using a reference genome-based approach.

The standard RNA-seq workflow for well-annotated genomes relies on mapping sequencing reads to a reference genome, then quantifying gene-level expression. This approach differs from de novo transcriptome assembly, which is necessary for organisms without reference genomes. For non-model organisms, transcriptome assembly and annotation require additional steps, including multi-BLAST annotation strategies to improve primary annotation quality. The workflow described here focuses on the reference-based approach, which is appropriate when a high-quality genome assembly and gene annotation are available.

## Galaxy Platform Overview

Galaxy is an open-source, web-based platform that provides access to hundreds of bioinformatics tools through a graphical interface. The platform supports persistent storage, data exchange, and documentation of intermediate results and analysis workflows. This design enables reproducibility because every analysis step is recorded and can be shared or rerun. The Galaxy Training Network provides accessible workflow training and analysis tutorials that cover the complete RNA-seq analysis path from raw data to biological interpretation.

The platform integrates tools for short-read alignment, transcript identification and quantification, and differential expression analysis. Workbenches such as Oqtans have demonstrated how Galaxy can support complete quantitative transcriptome analysis with customizable computational workflows and modular pipeline architecture. The Galaxy Toolshed allows users to install additional tools, extending the platform's capabilities beyond the default installation.

For researchers working with incompletely annotated species, Galaxy workflows can incorporate orthologue species information to increase the number of annotated miRNAs and avoid falsely annotated sequences. This flexibility makes Galaxy suitable for both model and non-model organisms, though the reference-based workflow described here requires a genome assembly and annotation.

## Data Inputs and Experimental Design

### Raw Sequencing Data Formats

RNA-seq experiments generate raw sequencing data in FASTQ format, which contains nucleotide sequences and quality scores for each read. These files are typically produced by Illumina sequencing platforms and can be large, often requiring substantial storage space. Raw data may be deposited in public repositories such as the NCBI Sequence Read Archive, which provides search systems and sequence resources for accessing published datasets. The NCBI maintains databases for sequence data, and researchers can download raw FASTQ files for reanalysis or use their own generated data.

### Experimental Design Considerations

Before beginning analysis, the experimental design must be evaluated. Biological replication is essential for reliable differential expression results. Studies using small numbers of biological replicates are unlikely to produce replicable results, and the combination of population heterogeneity and underpowered cohort sizes affects the replicability of RNA-seq research. Differential expression and enrichment analysis results from underpowered experiments are unlikely to replicate well, though some datasets achieve high median precision despite low recall and replicability for cohorts with more than five replicates.

The number of replicates required depends on the biological variability of the system under study, the magnitude of expression changes expected, and the desired statistical power. Researchers constrained by small cohort sizes can use bootstrapping procedures to estimate the expected performance regime of their datasets. This approach correlates strongly with observed replicability and precision metrics and provides practical guidance for experimental planning.

### Reference Genome and Annotation Requirements

The reference-based workflow requires a genome assembly in FASTA format and gene annotation in GTF or GFF format. These files are available from public databases for model organisms and many agricultural species. The quality of the reference genome directly affects mapping rates and quantification accuracy. For organisms with incomplete annotations, additional steps may be needed to improve gene models before quantification.

## Quality Control of Raw Reads

### Initial Quality Assessment

Quality control begins with examining the raw FASTQ files to identify problems such as adapter contamination, low-quality bases, and GC bias. The Galaxy platform provides tools for generating quality reports that summarize per-base quality scores, per-sequence quality scores, GC content, and adapter content. These reports allow researchers to make informed decisions about trimming parameters and whether any samples should be excluded from downstream analysis.

### Adapter Trimming and Read Filtering

Adapter sequences are commonly present in RNA-seq libraries, particularly when fragment sizes are shorter than the read length. Trimming removes adapter sequences and low-quality bases from the ends of reads. The choice of trimming parameters affects the number of reads retained and the quality of alignments. Overly aggressive trimming can remove biological sequence, while insufficient trimming leaves adapter contamination that interferes with alignment.

Quality filtering removes reads that fail to meet minimum quality thresholds. The specific thresholds depend on the sequencing platform and the downstream analysis requirements. After trimming and filtering, the quality control step should be repeated to confirm that the processed reads meet acceptable quality standards.

### Quality Control Records

Documenting quality control metrics for each sample provides a record of data quality that supports interpretation of downstream results. Key metrics include total read count, percentage of reads retained after trimming, mean quality scores, and GC content. These records allow researchers to identify problematic samples and to report data quality in publications.

## Read Alignment to Reference Genome

### Alignment Tool Selection

The choice of alignment tool affects mapping accuracy, speed, and the format of output files. HISAT2 is a widely used splice-aware aligner that can efficiently map RNA-seq reads to a reference genome. It is part of standard RNA-seq pipelines and has been incorporated into automated analysis platforms. The Galaxy platform provides HISAT2 and other aligners through its tool interface, allowing researchers to select the appropriate tool for their data.

Splice-aware alignment is essential for RNA-seq data because reads spanning exon-exon junctions do not align contiguously to the genome. Aligners must identify splice junctions and map reads across them correctly. The accuracy of splice junction detection affects quantification of genes with multiple isoforms and the identification of differentially spliced transcripts.

### Alignment Parameters

Alignment parameters include the maximum number of mismatches allowed, the number of alignments reported per read, and settings for handling multi-mapping reads. Reads that map to multiple locations in the genome present a challenge for quantification because their origin is ambiguous. Common approaches include discarding multi-mapping reads, distributing them proportionally, or assigning them to a single location based on alignment score.

The choice of parameters should be documented and applied consistently across all samples in an experiment. Changes in alignment parameters between samples can introduce systematic bias that affects differential expression results.

### Alignment Quality Assessment

After alignment, the percentage of reads that map to the reference genome provides an initial indication of data quality. Low mapping rates may indicate sample contamination, adapter contamination, or problems with the reference genome. The distribution of reads across genomic features, such as exons, introns, and intergenic regions, provides additional quality information. High rates of reads mapping to introns may indicate genomic DNA contamination, while high rates of intergenic mapping may indicate problems with the annotation.

## Quantification of Gene Expression

### Counting Reads in Genomic Features

FeatureCounts is a widely used tool for assigning aligned reads to genomic features and generating count tables. It takes aligned reads in BAM format and a gene annotation file as input, then produces a count matrix with genes as rows and samples as columns. The count matrix serves as the input for differential expression analysis.

The choice of counting parameters affects the results. Parameters include whether to count reads that map to multiple locations, how to handle reads that overlap multiple features, and whether to count reads that span exon-exon junctions. These decisions should be based on the experimental questions and the characteristics of the annotation.

### Transcript Assembly and Quantification Alternatives

For some analyses, transcript-level quantification is preferred over gene-level counting. Tools such as StringTie assemble transcripts from aligned reads and estimate transcript abundances. This approach can identify novel isoforms and provide isoform-level expression estimates. However, transcript-level analysis is more complex and may require additional validation.

The choice between gene-level counting and transcript-level assembly depends on the biological questions. Gene-level analysis is appropriate for identifying differentially expressed genes, while transcript-level analysis is needed for studying alternative splicing and isoform switching.

### Normalization Considerations

Raw read counts are not directly comparable across samples because of differences in sequencing depth and library composition. Normalization methods adjust counts to account for these technical factors. The choice of normalization method affects differential expression results, particularly for genes with extreme expression levels or for experiments with large differences in library composition.

For datasets with limited replicates, an alternative workflow based on transcripts per million may be used. This approach normalizes counts by transcript length and sequencing depth, providing a measure of relative abundance that can be compared across samples.

## Differential Expression Analysis

### Statistical Methods

DESeq2 is a widely used method for differential expression analysis that models count data with a negative binomial distribution. It estimates dispersion and tests for differences in expression between conditions while accounting for biological variability. DESeq2 is integrated into Galaxy and is part of standard RNA-seq analysis pipelines.

The statistical analysis produces a table of results with log2 fold changes, p-values, and adjusted p-values for each gene. The adjusted p-values account for multiple testing, which is necessary because thousands of genes are tested simultaneously. The choice of significance thresholds affects the number of genes identified as differentially expressed and the false discovery rate.

### Interpretation of Results

Differentially expressed genes are typically defined by a combination of fold change and adjusted p-value thresholds. The specific thresholds depend on the experimental context and the goals of the study. Genes with large fold changes and small adjusted p-values are the most confidently identified as differentially expressed.

The biological interpretation of differential expression results requires connecting the identified genes to biological processes and pathways. Functional enrichment analysis using Gene Ontology, Kyoto Encyclopedia of Genes and Genomes, Reactome, and Gene Set Enrichment Analysis can identify biological processes that are overrepresented among differentially expressed genes. These analyses provide context for interpreting the biological significance of expression changes.

### Validation of Results

Validation of differential expression results using an independent method, such as quantitative real-time reverse transcription polymerase chain reaction, provides confidence in the findings. RNA-seq and RT-qPCR methods reveal similar directions of gene expression changes, but demonstrate differences in effect size and sensitivity. This emphasizes the need for a combined use of both approaches when precise quantification is required.

## At a Glance

| Workflow Step | Primary Tool Options | Key Parameters and Decisions | Output |
| --- | --- | --- | --- |
| Quality Control | FastQC, Trimmomatic | Adapter trimming, quality thresholds, read filtering | Quality reports, trimmed FASTQ files |
| Read Alignment | HISAT2 | Splice-aware mapping, mismatch allowance, multi-mapping handling | BAM alignment files |
| Quantification | FeatureCounts, StringTie | Feature type, multi-mapping reads, strand specificity | Count matrix or transcript abundances |
| Differential Expression | DESeq2 | Normalization method, significance thresholds, dispersion estimation | Results table with fold changes and adjusted p-values |
| Functional Enrichment | Gene Ontology, KEGG, Reactome, GSEA | Reference gene sets, significance thresholds | Enriched biological processes and pathways |

## Practical Workflow Implementation

### Step 1: Data Preparation

Upload raw FASTQ files to Galaxy and organize them into a history. Create a history for each analysis project to maintain organization and reproducibility. Verify that all files are complete and correctly labeled before proceeding.

### Step 2: Quality Control

Run quality assessment on all raw FASTQ files. Review the quality reports and identify any samples with unusual patterns. Document the quality metrics for each sample in a laboratory notebook or electronic record.

### Step 3: Trimming and Filtering

Based on the quality reports, select trimming parameters appropriate for the data. Run adapter trimming and quality filtering on all samples. Repeat quality assessment on the trimmed reads to confirm improvement.

### Step 4: Alignment

Select the reference genome and annotation files appropriate for the organism under study. Run splice-aware alignment on all trimmed FASTQ files. Assess alignment rates and investigate any samples with unusually low mapping percentages.

### Step 5: Quantification

Run feature counting on all aligned samples using the gene annotation file. Review the count matrix for any samples with unusually low total counts or other anomalies.

### Step 6: Differential Expression Analysis

Run differential expression analysis comparing the conditions of interest. Review the results table and identify genes meeting the significance thresholds. Generate visualizations such as MA plots and heat maps to explore the results.

### Step 7: Functional Enrichment

Run functional enrichment analysis on the list of differentially expressed genes. Interpret the enriched biological processes in the context of the experimental system.

### Step 8: Documentation

Record all tool versions, parameters, and settings used in the analysis. Save the Galaxy workflow so that the analysis can be rerun or shared with collaborators. Export the final results tables for use in publications.

## Records and Measurements

### Essential Records for Reproducibility

Reproducibility requires complete documentation of the analysis. Galaxy automatically records the history of tools and parameters used, providing a built-in record of the analysis. Exporting the workflow and history creates a permanent record that can be shared with collaborators or included in supplementary materials.

Key records include the version of each tool used, the exact parameters selected, the reference genome and annotation versions, and the input file names. Changes in any of these components can affect results, so accurate records are essential for reproducing the analysis.

### Measurements to Track

The following measurements should be tracked for each sample throughout the workflow:

- Total number of raw reads
- Percentage of reads retained after trimming
- Percentage of reads aligned to the reference genome
- Percentage of reads assigned to genes
- Total number of genes detected
- Distribution of read counts across samples

These measurements provide quality indicators and help identify problematic samples. Unusual values for any measurement should be investigated before proceeding with downstream analysis.

## Common Failure Patterns

### Low Mapping Rates

Low percentages of reads mapping to the reference genome can result from sample contamination, adapter contamination, or mismatches between the sample and reference genome. Investigate the quality reports and check for the presence of unexpected sequences. If the organism is closely related to the reference but not identical, consider using a more appropriate reference or adjusting alignment parameters.

### Batch Effects

Samples processed in different batches can show systematic differences in expression that are unrelated to biological conditions. Batch effects can obscure true biological differences or create spurious differences. Including batch information in the statistical model can account for these effects when the experimental design is balanced.

### Outlier Samples

Individual samples that differ substantially from others in the same condition can distort differential expression results. Examine quality metrics and count distributions to identify outliers. Consider whether technical problems explain the outlier pattern and whether the sample should be excluded from analysis.

### Overdispersion

RNA-seq count data often show more variability than expected under a simple Poisson model. DESeq2 accounts for this overdispersion by estimating dispersion parameters. Samples with unusually high variability may indicate technical problems or contamination.

### Underpowered Experiments

Experiments with few biological replicates have limited power to detect differential expression. Results from underpowered experiments are unlikely to replicate well, and the specific genes identified may not be reproducible. Researchers should interpret results from small cohorts with caution and consider validation with additional samples or independent methods.

## Limitations and Interpretation Boundaries

### Technical Limitations

RNA-seq analysis provides relative measures of gene expression, not absolute quantities. Comparisons between genes within a sample are limited by differences in transcript length and sequencing efficiency. Comparisons between experiments require careful normalization and should be interpreted with caution.

The reference-based approach depends on the quality of the genome assembly and annotation. Genes that are absent from the annotation will not be quantified. Novel transcripts and isoforms that are not represented in the annotation will be missed.

### Biological Interpretation Limits

Differential expression analysis identifies changes in transcript abundance but does not directly measure protein levels or activity. Post-transcriptional regulation can decouple mRNA levels from protein levels. Changes in transcript abundance may reflect changes in transcription, mRNA stability, or both.

The biological significance of differential expression depends on the experimental context. A gene with a small fold change may be biologically important if it is a transcription factor or signaling molecule, while a gene with a large fold change may have limited biological impact if it is not functionally relevant to the system under study.

### Replicability Concerns

RNA-seq studies have demonstrated irreproducibility, particularly when experiments are underpowered. The combination of population heterogeneity and small cohort sizes affects the replicability of differential expression and enrichment analysis results. Researchers should design experiments with adequate replication and interpret results with appropriate caution.

## Safety and Regulatory Context

### Data Management

RNA-seq data may include sensitive information, particularly when derived from human subjects. Researchers must comply with institutional and regulatory requirements for data storage, sharing, and privacy. Public repositories such as NCBI provide mechanisms for controlled access to sensitive data.

### Computational Resource Use

The Galaxy public server provides computational resources for analysis, but large datasets may require substantial processing time. Researchers should be mindful of resource limits and plan analyses accordingly. For very large datasets, local installation of Galaxy or use of cloud resources may be appropriate.

### Publication Requirements

Many journals require that raw sequencing data be deposited in public repositories and that analysis workflows be documented. Galaxy workflows provide a mechanism for sharing complete analysis pipelines, supporting the reproducibility requirements of scientific publication.

## Professional Escalation Criteria

Researchers should seek assistance from bioinformatics specialists or experienced colleagues when encountering the following situations:

- Mapping rates are consistently below expected values for the organism and library type
- Quality metrics indicate severe adapter contamination or other technical problems
- Differential expression results are highly unstable across analysis parameter choices
- The analysis requires tools or features not available in the Galaxy public server
- The experimental design requires advanced statistical modeling beyond standard differential expression analysis
- Results will be used for clinical decisions or regulatory submissions

## Decision Framework for Tool Selection and Parameter Choices in Galaxy RNA-seq Analysis

Selecting the correct tools and parameters in Galaxy is not a single decision but a sequence of choices that each affect the biological conclusions drawn from the data. Researchers without programming skills often accept default settings without understanding the consequences, which can lead to results that are technically valid but biologically misleading. This section provides a practical decision framework for choosing among alignment tools, quantification methods, and statistical approaches, along with a structured record system for documenting those choices.

### Alignment Tool Selection Framework

The alignment step converts raw sequencing reads into genomic coordinates, and the choice of aligner determines how splice junctions are handled, how multi-mapping reads are resolved, and how alignment errors are scored. HISAT2 is the most commonly used splice-aware aligner in Galaxy workflows and is included in standard RNA-seq pipelines. It uses a hierarchical indexing strategy that balances speed and memory usage, making it suitable for the Galaxy public server where computational resources are shared among many users.

For most reference-based RNA-seq experiments, HISAT2 is the appropriate default choice. The decision framework begins with three questions about the experimental system. First, is the reference genome well annotated with known splice junctions? If yes, HISAT2 can use the annotation to improve junction detection. Second, are the reads of standard length, typically 50 to 150 base pairs? HISAT2 performs well across this range. Third, is the organism a common model species with a high-quality reference? If the answer to all three questions is yes, proceed with HISAT2 using default parameters.

When the reference genome is fragmented or contains many assembly errors, alternative aligners that are more tolerant of mismatches may be appropriate. When the organism is closely related to the reference species but not identical, alignment parameters should be adjusted to allow more mismatches. The Galaxy interface exposes these parameters, but researchers should understand that increasing mismatch tolerance also increases the rate of spurious alignments.

The decision to use a splice-aware aligner versus a simple unspliced aligner is not optional for RNA-seq data. Reads that span exon-exon junctions do not align contiguously to the genome, and only splice-aware aligners can map them correctly. Using an unspliced aligner would result in substantial read loss and biased quantification of multi-exon genes.

### Quantification Method Selection

After alignment, the researcher must choose between gene-level counting and transcript-level assembly. FeatureCounts is the standard tool for gene-level counting in Galaxy and assigns aligned reads to genomic features defined in the annotation file. StringTie performs transcript assembly and abundance estimation, providing isoform-level information.

The decision between these approaches depends on the biological question. Gene-level counting with FeatureCounts is appropriate when the goal is to identify differentially expressed genes. This approach is simpler, more robust, and requires fewer computational resources. Transcript-level analysis with StringTie is necessary when studying alternative splicing, isoform switching, or novel transcript discovery.

For gene-level counting, the key parameters are the feature type and the handling of multi-mapping reads. The feature type parameter specifies which annotation features to count, typically exon or gene. Counting at the gene level sums reads across all exons of a gene, while counting at the exon level provides exon-specific counts that can be used for splicing analysis. The multi-mapping parameter determines whether reads that align to multiple genomic locations are counted, counted fractionally, or discarded.

The decision framework for quantification includes a check on annotation completeness. If the annotation is incomplete, gene-level counting will underestimate expression for genes with missing exons. In this case, transcript assembly with StringTie can identify novel transcripts, but the results require additional validation before they can be used for differential expression analysis.

### Statistical Method Selection

DESeq2 is the standard differential expression tool in Galaxy and models count data with a negative binomial distribution. It estimates dispersion for each gene and shrinks these estimates toward a common value, which improves stability for genes with low read counts. The choice of statistical method is less flexible in Galaxy than the choice of aligner or quantifier, but the parameters within DESeq2 require careful consideration.

The key parameters in DESeq2 are the significance threshold, the fold change threshold, and the handling of genes with low counts. The adjusted p-value threshold controls the false discovery rate and is typically set to 0.05 or 0.01. The fold change threshold specifies the minimum magnitude of expression change required for a gene to be reported as differentially expressed. The default fold change threshold is often zero, meaning any gene with a statistically significant difference is reported regardless of effect size.

The decision framework for statistical analysis includes a check on the experimental design. DESeq2 can incorporate multiple factors in the model, such as treatment and batch. Including batch as a factor in the model accounts for systematic technical variation when the experimental design is balanced. The design formula must be specified correctly, and researchers should verify that the factor levels are correctly assigned to samples.

### Parameter Documentation System

Reproducibility in Galaxy is supported by the automatic recording of tool versions and parameters in the analysis history. However, the history does not capture the reasoning behind parameter choices. A structured documentation system that records the decision framework for each analysis provides context that is essential for interpreting results and for troubleshooting when results are unexpected.

The documentation system should include a parameter decision log with entries for each tool in the workflow. Each entry records the tool name, version, all parameters changed from defaults, the rationale for each change, and the date of the analysis. This log can be maintained in a spreadsheet or laboratory notebook and should be updated whenever the workflow is modified.

The parameter decision log serves multiple purposes. It provides a record for publications that require detailed methods descriptions. It enables comparison between analyses when parameters are changed. It supports troubleshooting by identifying which parameter changes affected the results. And it provides a training resource for new researchers who need to understand why specific choices were made.

### Comparison of Tool Outputs

A practical approach to validating tool choices is to compare outputs from different tools or parameter settings on a subset of data. This comparison can identify whether results are robust to reasonable parameter variations or whether they depend critically on specific choices.

For alignment, the comparison metric is the mapping rate and the distribution of reads across genomic features. If two aligners produce similar mapping rates and similar distributions, the choice of aligner is unlikely to affect downstream results. If the mapping rates differ substantially, the aligner choice is consequential and should be investigated.

For quantification, the comparison metric is the correlation between count matrices. High correlation between FeatureCounts and StringTie outputs for the same samples indicates that gene-level results are robust to the quantification method. Low correlation suggests that the choice of quantification method affects the results, and the source of the discrepancy should be investigated.

For differential expression, the comparison metric is the overlap between significant gene lists. If two statistical methods or parameter settings produce largely overlapping lists of differentially expressed genes, the results are robust. If the overlap is low, the results depend on the specific choices and should be interpreted with caution.

### Decision Points for Non-Model Organisms

Researchers working with non-model organisms face additional decision points that do not arise for well-annotated model species. The reference genome may be fragmented, the annotation may be incomplete, or no reference genome may be available at all.

When a reference genome is available but the annotation is incomplete, the decision framework includes an assessment of annotation quality. The percentage of reads that map to annotated exons provides a measure of annotation completeness. Low exon mapping rates indicate that many transcripts are not represented in the annotation, and gene-level counting will underestimate expression for these genes.

For incompletely annotated species, the decision framework may include orthologue-based annotation improvement. Galaxy workflows can incorporate orthologue species information to increase the number of annotated genes and avoid falsely annotated sequences. This approach was demonstrated for miRNA analysis in species with incomplete annotations, where orthologue information increased the total number of annotated miRNAs.

When no reference genome is available, the reference-based workflow described here cannot be used. De novo transcriptome assembly is required, which involves additional steps for transcript assembly and annotation. The Galaxy platform supports this approach, but the workflow is more complex and requires additional expertise.

### Resource Allocation Decisions

The Galaxy public server provides shared computational resources, and large datasets can require substantial processing time. The decision framework includes an assessment of computational requirements before beginning the analysis.

The number of samples, read depth, and genome size all affect processing time. A typical RNA-seq experiment with 12 samples and 30 million reads per sample requires significant processing time on the Galaxy public server. Researchers should plan for this time and consider whether the analysis can be divided into smaller batches.

For very large datasets, local installation of Galaxy or use of cloud resources may be appropriate. The Galaxy platform supports both options, and the choice depends on the available local infrastructure and the scale of the analysis. The decision framework should include an estimate of processing time and a plan for managing large analyses.

### Validation of Parameter Choices

The decision framework includes a validation step that checks whether the chosen parameters produce biologically sensible results. This validation is distinct from the statistical validation of differential expression results and focuses on the technical quality of the analysis.

The first validation check is the distribution of adjusted p-values. Under the null hypothesis of no differential expression, adjusted p-values should be approximately uniformly distributed. A large spike near zero may indicate that the statistical model is not appropriate for the data, possibly because of unaccounted technical variation.

The second validation check is the relationship between mean expression and fold change. In most experiments, genes with low expression show more variable fold changes than genes with high expression. DESeq2 accounts for this relationship through dispersion estimation, but extreme patterns may indicate problems with the data or the model.

The third validation check is the consistency of results across parameter settings. If changing the fold change threshold from 0 to 1 dramatically changes the list of differentially expressed genes, the results are sensitive to this parameter and should be interpreted with caution.

### Record System for Analysis Decisions

The record system for analysis decisions extends beyond the parameter decision log to include the complete analysis trail. This system captures the sequence of decisions, the data used at each step, and the rationale for each choice.

The analysis trail record includes the following components:

- The raw data file names and their source
- The reference genome and annotation versions
- The quality control metrics for each sample
- The alignment tool and parameters
- The quantification tool and parameters
- The differential expression tool and parameters
- The significance thresholds and the number of genes identified at each threshold
- The functional enrichment results and the reference gene sets used

This record system is maintained throughout the analysis and updated at each step. The Galaxy history provides the technical record of tools and parameters, while the analysis trail record provides the biological context and decision rationale.

### Common Decision Errors

Several common decision errors lead to problematic results in Galaxy RNA-seq analysis. Recognizing these errors is the first step in avoiding them.

The first error is accepting default parameters without understanding their implications. Default parameters are designed for typical experiments, but typical is not universal. The decision framework should be applied to each experiment to determine whether defaults are appropriate.

The second error is changing parameters without documenting the change. This error undermines reproducibility and makes troubleshooting difficult. The parameter decision log should be updated whenever any parameter is changed.

The third error is using the same parameters for all samples without checking whether individual samples require different treatment. Outlier samples may require additional trimming or filtering, and applying uniform parameters to all samples can obscure sample-specific problems.

The fourth error is interpreting results without considering the limitations of the analysis. The decision framework should include an assessment of what the analysis can and cannot detect, given the experimental design and data quality.

### Escalation Criteria for Decision Support

Some decisions require expertise beyond what the decision framework provides. The following situations warrant escalation to a bioinformatics specialist:

- The reference genome is highly fragmented or contains many assembly errors
- The annotation is incomplete and orthologue-based improvement is needed
- The experimental design includes complex factors that require advanced statistical modeling
- The results are highly sensitive to parameter choices and no clear biological rationale exists for the preferred settings
- The analysis requires tools not available in the Galaxy public server
- The results will be used for clinical decisions or regulatory submissions

In these situations, the researcher should document the decision points and the data that inform them, then seek assistance from someone with specialized expertise. The documentation provides the context needed for the specialist to provide useful guidance.

### Decision Framework Summary Table

The following table summarizes the decision framework for each major step in the workflow:

| Decision Point | Primary Options | Key Considerations | Default Recommendation |
| --- | --- | --- | --- |
| Alignment tool | HISAT2, other splice-aware aligners | Genome quality, read length, organism relatedness to reference | HISAT2 with default parameters |
| Quantification method | FeatureCounts, StringTie | Biological question, annotation completeness | FeatureCounts for gene-level analysis |
| Multi-mapping reads | Count, fractional, discard | Genome complexity, read length | Discard for most analyses |
| Normalization | DESeq2 internal, TPM | Replicate number, library composition | DESeq2 internal normalization |
| Significance threshold | Adjusted p-value 0.05 or 0.01 | Expected number of true positives | Adjusted p-value 0.05 |
| Fold change threshold | 0, 1, 2 | Biological relevance, effect size expectations | 0 for discovery, 1 for focused analysis |
| Batch handling | Include as factor, ignore | Experimental design balance | Include as factor when balanced |

### Applying the Decision Framework

The decision framework is applied at the beginning of each analysis and revisited at each step. The framework is not a one-time assessment but an ongoing process that responds to the data as it is generated.

At the start of the analysis, the researcher assesses the experimental design, the reference genome quality, and the expected data characteristics. This assessment informs the initial tool and parameter choices. As quality control results become available, the researcher may adjust trimming parameters. As alignment results become available, the researcher may adjust alignment parameters or investigate problematic samples. As differential expression results become available, the researcher may adjust significance thresholds based on the distribution of results.

The decision framework is designed to be practical and actionable for researchers without programming skills. It provides a structured approach to making choices that are often made arbitrarily or by default. By applying the framework and documenting the decisions, researchers can produce reproducible analyses that support reliable biological conclusions.

The framework also supports the comparison of results across studies. When the decision framework is documented, other researchers can understand the choices that were made and assess whether those choices are appropriate for their own data. This transparency is essential for the cumulative progress of transcriptome research, where the replicability of results depends on the consistency of analysis decisions.

## Frequently Asked Questions

### What is the difference between raw counts and normalized expression values?

Raw counts are the number of sequencing reads assigned to each gene. These values are not directly comparable across samples because of differences in sequencing depth. Normalized expression values adjust for sequencing depth and other technical factors, allowing comparisons between samples. Differential expression analysis uses statistical models that account for these technical factors while testing for biological differences.

### How many biological replicates are needed for reliable differential expression results?

The number of replicates needed depends on the biological variability of the system and the magnitude of expression changes. Studies with small numbers of replicates are unlikely to produce replicable results. More than five replicates per condition generally improve replicability, though the specific requirements depend on the experimental system. Researchers can use bootstrapping procedures to estimate the expected performance of their experimental design.

### What should I do if my mapping rate is low?

Low mapping rates can result from sample contamination, adapter contamination, or mismatches between the sample and reference genome. Check the quality reports for evidence of adapter contamination or low-quality reads. Verify that the reference genome matches the organism under study. Consider whether the sample may be contaminated with DNA from another organism.

### Can I use Galaxy for non-model organisms without a reference genome?

Galaxy supports de novo transcriptome assembly for organisms without reference genomes. This approach requires additional steps for transcript assembly and annotation. For incompletely annotated species, orthologue species information can improve annotation. The workflow is more complex than the reference-based approach and may require additional expertise.

### What is the difference between gene-level and transcript-level analysis?

Gene-level analysis counts reads assigned to all transcripts of a gene, providing a measure of total gene expression. Transcript-level analysis estimates the abundance of individual isoforms, allowing detection of isoform switching and alternative splicing. Transcript-level analysis is more complex and requires additional computational steps.

### How do I choose significance thresholds for differential expression?

Significance thresholds should be based on the goals of the study and the expected number of true positives. Adjusted p-values account for multiple testing and control the false discovery rate. Common thresholds include adjusted p-values below 0.05 or 0.01, often combined with a minimum fold change. The specific thresholds should be justified in the context of the experimental questions.

### What causes batch effects and how can I account for them?

Batch effects are systematic differences between samples processed at different times or in different groups. They can result from reagent lot changes, instrument differences, or processing date effects. Including batch information in the statistical model can account for these effects when the experimental design is balanced. Visualizing the data with principal component analysis can help identify batch effects.

### How should I validate my differential expression results?

Validation with an independent method, such as quantitative real-time reverse transcription polymerase chain reaction, provides confidence in the findings. RNA-seq and RT-qPCR methods show similar directions of gene expression changes but differ in effect size and sensitivity. Validation is particularly important for genes that will be studied further or used for hypothesis generation.

## Related Bioinformatics Guides

- [RNA-Seq Data Analysis in Galaxy: A User-Friendly Platform](/knowledge/bioinformatics/rna-seq-data-analysis-in-galaxy-a-user-friendly-platform)
- [RNA-Seq Data Analysis Workflow: From Raw Reads to Insights](/knowledge/bioinformatics/rna-seq-data-analysis-workflow-from-raw-reads-to-insights)
- [RNA Sequencing Data Analysis: From Raw Reads to Differential Expression](/knowledge/bioinformatics/rna-sequencing-data-analysis-from-raw-reads-to-differential-expression)
- [Proteomics Data Analysis in R: A Practical Workflow for Differential Expression and Visualization](/knowledge/bioinformatics/proteomics-data-analysis-in-r-a-practical-workflow-for-differential-expression-and-visualization)
- [Metabolomics Data Analysis Workflow: From Raw Data to Biological Insight](/knowledge/bioinformatics/metabolomics-data-analysis-workflow-from-raw-data-to-biological-insight)

## Related Clinical & Scientific Guides

* [A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data](/knowledge/bioinformatics/a-practical-guide-to-detecting-antimicrobial-resistance-genes-in-shotgun-metagenomic-data)
* [Computational Immunology: Modeling the Immune System](/knowledge/bioinformatics/computational-immunology-modeling-the-immune-system)
* [How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices](/knowledge/bioinformatics/how-to-set-hard-filters-for-germline-variant-calling-a-practical-guide-to-gatk-best-practices)


## References and Further Reading

- [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information.
- [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute.
- [Bioconductor](https://bioconductor.org/). Bioconductor Project.
- [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.
- [nf-core Documentation](https://nf-co.re/docs). nf-core.
- [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries.
- [Reference-Based Gene Expression Analysis Using Galaxy Server.](https://pubmed.ncbi.nlm.nih.gov/40067633). Methods in molecular biology (Clifton, N.J.), 2025.
- [Barley (Hordeum Vulgare) Anther and Meiocyte RNA Sequencing: Mapping Sequencing Reads and Downstream Data Analyses.](https://pubmed.ncbi.nlm.nih.gov/35461459). Methods in molecular biology (Clifton, N.J.), 2022.
- [Small RNA-seq analysis of single porcine blastocysts revealed that maternal estradiol-17beta exposure does not affect miRNA isoform (isomiR) expression.](https://pubmed.ncbi.nlm.nih.gov/30081835). BMC genomics, 2018.
- [Galaxy and MEAN Stack to Create a User-Friendly Workflow for the Rational Optimization of Cancer Chemotherapy.](https://pubmed.ncbi.nlm.nih.gov/33679888). Frontiers in genetics, 2021.
- [Oqtans: the RNA-seq workbench in the cloud for complete and reproducible quantitative transcriptome analysis.](https://pubmed.ncbi.nlm.nih.gov/24413671). Bioinformatics (Oxford, England), 2014.
- [Virus-independent and common transcriptome responses of leafhopper vectors feeding on maize infected with semi-persistently and persistent propagatively transmitted viruses.](https://pubmed.ncbi.nlm.nih.gov/24524215). BMC genomics, 2014.
- [Studying Oogenesis in a Non-model Organism Using Transcriptomics: Assembling, Annotating, and Analyzing Your Data.](https://pubmed.ncbi.nlm.nih.gov/27557578). Methods in molecular biology (Clifton, N.J.), 2016.
- [Development and validation of the PipeSeq program for RNA-seq data analysis in the Chlamydomonas reinhardtii as a model.](https://doi.org/10.18699/vjgb-26-34). 2026.
- [RAGER: A user-friendly computational platform for integrated analysis of RNA-Seq and ATAC-seq data.](https://doi.org/10.1371/journal.pone.0349941). 2026.
- [Metatranscriptomic Reanalysis of Alzheimer's Brains Identifies Low-Biomass Microbial Signals Including Enrichment of &lt,i&gt,Acinetobacter radioresistens&lt,/i&gt,.](https://doi.org/10.3390/ijms27083430). 2026.
- [THRAISE: An automated and reproducible web platform for RNA-seq analysis.](https://doi.org/10.1016/j.mocell.2025.100299). 2025.
- [CRESCENT, a comprehensive RNA-Seq expression, splicing, and coding/non-coding element network tool.](https://doi.org/10.1186/s12859-026-06368-5). 2026.
- [Cross-species transcriptomic integration reveals a MIRO1-mediated macrophage-T cell axis in glioma.](https://doi.org/10.26508/lsa.202603749). 2026.
- [Transcriptomic rewiring of the JAK-STAT pathway in circulating CD4&lt,sup&gt,+&lt,/sup&gt,CLA&lt,sup&gt,+&lt,/sup&gt, and CD4&lt,sup&gt,+&lt,/sup&gt, naïve T cells from patients with atopic dermatitis and psoriasis.](https://doi.org/10.3389/fimmu.2026.1782684). 2026.
- [RNAflow: An Effective and Simple RNA-Seq Differential Gene Expression Pipeline Using Nextflow](https://doi.org/10.3390/genes11121487). Genes, 2020.
- [Replicability of bulk RNA-Seq differential expression and enrichment analysis results for small cohort sizes](https://doi.org/10.1371/journal.pcbi.1011630). PLoS Comput. Biol., 2025.
- [RNA-Seq Data Analysis for Differential Gene Expression Using HISAT2-StringTie-Ballgown Pipeline.](https://doi.org/10.1007/978-1-0716-3886-6_5). Methods in molecular biology, 2024.
- [Transcriptome meta-analysis of valproic acid exposure in human embryonic stem cells](https://doi.org/10.1016/j.euroneuro.2022.04.008). European Neuropsychopharmacology, 2022.

> This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.