How to Perform Gene Ontology (GO) Enrichment Analysis in R: A Step-by-Step Tutorial Using clusterProfiler

By Dr. Zubair Khalid, DVM, MS, PhD ·

How to Perform Gene Ontology (GO) Enrichment Analysis in R: A Step-by-Step Tutorial Using clusterProfiler

Key Takeaways

  • Gene Ontology (GO) enrichment analysis identifies functional categories over-represented in a gene list, answering "which biological processes, molecular functions, or cellular components are disproportionately represented among differentially expressed genes?" This is achieved by comparing the proportion of genes in a specific GO term within the experimental gene list against the proportion in a defined background set of all tested genes.
  • The clusterProfiler package in R, distributed via Bioconductor, is a robust tool for GO enrichment, supporting over-representation analysis (ORA) with enrichGO and functional class scoring (FCS) with gseGO. ORA requires a binary gene list, while FCS utilizes a ranked list of genes, offering greater sensitivity for detecting subtle pathway changes.
  • Accurate definition of the "background gene set" is critical; it must comprise all genes that were subjected to differential expression testing, not the entire genome, to avoid inflating significance. For RNA-seq, this typically includes genes with sufficient expression, while for microarrays, it's all genes on the array.
  • Gene identifier mapping (e.g., from gene symbols to Entrez IDs) using functions like bitr is essential for clusterProfiler to interface with organism-specific annotation packages (e.g., org.Hs.eg.db for humans). Failure to map correctly or using incompatible identifier types is a common cause of empty results.
  • Visualizations such as dot plots (dotplot) and term-gene network plots (cnetplot) are crucial for interpreting enrichment results, illustrating term significance, gene ratios, and gene-term relationships, while redundancy reduction techniques (e.g., simplify) are necessary to manage the hierarchical nature of GO and avoid obscuring biological signals.
  • Reproducibility is paramount; record all parameters, including annotation package versions, significance thresholds (e.g., adjusted p-value < 0.05, absolute log2 fold change > 1), and multiple testing correction methods (e.g., Benjamini-Hochberg), and use sessionInfo() to document the computational environment.

Gene Ontology (GO) enrichment analysis answers a specific question after differential expression testing: among the genes that changed in your experiment, which functional categories appear more often than expected by chance? This tutorial walks through the complete process using the clusterProfiler package in R, from preparing your gene list to interpreting and visualizing enrichment results. You will learn how to choose the correct background set, run over-representation analysis, control for multiple testing, and produce publication-ready figures. The workflow applies to RNA-seq differential expression output, microarray results, and any other source of a ranked or selected gene list.

What GO Enrichment Analysis Does and Why It Matters

High-throughput sequencing routinely produces lists of differentially expressed genes (DEGs). A typical RNA-seq experiment comparing two conditions may yield hundreds or thousands of DEGs, and reading through that list gene by gene does not reveal the underlying biology. GO enrichment analysis groups those genes into functional categories based on three structured ontologies: biological process, molecular function, and cellular component. The analysis then asks whether certain categories contain more genes from your list than would be expected given the total set of genes measured in your experiment.

The practical value of this approach is well established in biomedical research. A bench scientist tutorial on functional class enrichment analysis describes the core question directly: which functional groups of genes among the DEGs are meaningful underlying factors to the differential biological or biomedical conditions under investigation [<a href="#ref-1">1</a>]. The same tutorial notes that R is a robust platform for this work because it is accessible to general users and has well-developed packages for enrichment analysis, visualization, and knowledge database access [<a href="#ref-1">1</a>].

GO enrichment is not limited to human or mouse studies. Researchers have applied it to congenital diaphragmatic hernia-associated genes to identify distinct biological pathways linked to different clinical presentations, including retinoic acid signaling, myogenesis, and angiogenesis [<a href="#ref-2">2</a>]. Plant researchers have used transcriptome analysis combined with enrichment to reveal sugar metabolism regulatory networks in maize stalks under different planting densities [<a href="#ref-3">3</a>]. Fungal genomics has a dedicated R package, FunFEA, that supports GO, KEGG, and COG/KOG enrichment for both model organisms and novel genomes with eggNOG-mapper annotations [<a href="#ref-4">4</a>]. These examples show that the same analytical framework transfers across species and experimental designs.

The analysis pipeline described in this tutorial uses clusterProfiler, which is distributed through Bioconductor. Bioconductor provides official documentation for package installation, workflows, and reproducible genomic analysis [<a href="#ref-5">5</a>]. The Nature Protocols publication on using clusterProfiler to characterize multiomics data confirms that the package supports enrichment analysis across multiple data types beyond simple gene lists [<a href="#ref-6">6</a>].

Preparing Your Input Data

Required Data Format

clusterProfiler accepts a character vector of gene identifiers as the primary input for over-representation analysis. The most common identifier types are Entrez Gene IDs, Ensembl gene IDs, and official gene symbols. Entrez Gene IDs are often preferred because they map directly to the annotation databases that clusterProfiler uses internally.

The National Center for Biotechnology Information maintains the databases that provide these identifiers and their functional annotations [<a href="#ref-7">7</a>]. NCBI resources include the Gene database for gene-specific information, the Sequence Read Archive for raw sequencing data, and the Gene Expression Omnibus for processed expression data [<a href="#ref-7">7</a>]. When you download annotation files or use annotation packages, you are drawing on this infrastructure.

Your input vector should contain only the genes you consider significant. For a typical RNA-seq analysis, this means genes that passed your thresholds for adjusted p-value and log2 fold change. The choice of thresholds directly affects your results, so record them clearly. A common starting point is an adjusted p-value below 0.05 and an absolute log2 fold change above 1, but you should adjust these based on your experimental design and the number of DEGs you obtain.

The Background Gene Set

The background set is the list of all genes that were tested for differential expression, beyond those that passed your significance thresholds. This distinction is critical. If you use all genes in the genome as background when your experiment only measured a subset, you will overestimate enrichment significance for categories that happen to be well represented on your assay platform.

For RNA-seq data, the background should be all genes with sufficient expression to be tested. Many pipelines define this as genes with a minimum count across samples. For microarray data, the background is all genes represented on the array. clusterProfiler allows you to specify the background using the universe parameter in the enrichment functions.

The choice of background is one of the most consequential decisions in the entire workflow. A study on congenital diaphragmatic hernia genes used a comprehensive list of 218 CDH-associated genes identified from the literature and performed large-scale gene function analysis using GO to identify significantly enriched biological pathways and molecular functions [<a href="#ref-2">2</a>]. The background in that case was the set of all genes considered in the analysis, not the entire genome.

Annotation Packages and Organism Support

clusterProfiler relies on organism-level annotation packages to map gene identifiers to GO terms. These packages are available through Bioconductor for common model organisms. For humans, the package is org.Hs.eg.db. For mice, it is org.Mm.eg.db. Other organisms have similar packages following the naming convention org.XX.eg.db.

If you work with a less common organism, you may need to provide your own annotation. The FunFEA package for fungal genomics addresses this problem by generating background frequency models from publicly available annotations and processing eggNOG-mapper annotations for novel genomes [<a href="#ref-4">4</a>]. This approach allows functional enrichment analysis of species without existing annotation packages.

Installing and Loading Required Packages

Installing Bioconductor Packages

Bioconductor packages require installation through the BiocManager package instead of the standard install.packages function. The Bioconductor project provides official installation documentation and workflow guidance [<a href="#ref-5">5</a>]. To install BiocManager and then clusterProfiler, use the following commands in your R console:

if (!require("BiocManager", quietly = TRUE))
    install.packages("BiocManager")

BiocManager::install("clusterProfiler")

You will also need the annotation package for your organism. For human data:

BiocManager::install("org.Hs.eg.db")

For visualization, you may want additional packages such as enrichplot for plotting enrichment results and DOSE for Disease Ontology enrichment. The Nature Protocols publication on clusterProfiler describes its use for characterizing multiomics data, which includes visualization and interpretation steps beyond basic enrichment [<a href="#ref-6">6</a>].

Loading Packages in Each Session

R does not remember installed packages between sessions. You must load them each time you start a new R session. Create a script header that loads all required packages:

library(clusterProfiler)
library(org.Hs.eg.db)
library(enrichplot)
library(ggplot2)

The Carpentries lessons provide foundational training in programming practices that apply here, including organizing code into scripts and documenting your workflow [<a href="#ref-8">8</a>]. A reproducible analysis script should include package loading, data import, analysis steps, and output generation in a single file.

Running Over-Representation Analysis with clusterProfiler

The Core Function: enrichGO

The primary function for GO enrichment in clusterProfiler is enrichGO. This function performs over-representation analysis (ORA), which compares the proportion of genes in your input list that belong to a GO term against the proportion in the background set.

The basic call requires several arguments:

ego <- enrichGO(
    gene = deg_entrez,
    universe = all_tested_entrez,
    OrgDb = org.Hs.eg.db,
    keyType = "ENTREZID",
    ont = "BP",
    pAdjustMethod = "BH",
    pvalueCutoff = 0.05,
    qvalueCutoff = 0.2,
    readable = TRUE
)

The gene argument is your vector of significant genes. The universe argument is your background set. The OrgDb argument specifies the annotation package. The keyType argument tells the function what type of identifiers you provided. The ont argument selects which GO ontology to test: "BP" for biological process, "MF" for molecular function, or "CC" for cellular component. You can also use "ALL" to test all three simultaneously.

The pAdjustMethod argument controls how p-values are corrected for multiple testing. The Benjamini-Hochberg method, specified as "BH", is the standard choice. The pvalueCutoff and qvalueCutoff arguments set thresholds for significance. The readable argument converts Entrez IDs to gene symbols in the output for easier interpretation.

Converting Gene Identifiers

Before running enrichGO, you must convert your gene identifiers to the type expected by the annotation package. If your differential expression results use Ensembl IDs or gene symbols, use the bitr function to convert them:

entrez_ids <- bitr(
    deg_symbols,
    fromType = "SYMBOL",
    toType = "ENTREZID",
    OrgDb = org.Hs.eg.db
)

This function maps identifiers using the annotation database and returns a data frame with both the original and converted identifiers. Some genes may not map successfully, particularly if they are non-coding RNAs or have been retired from the database. The unmapped genes are excluded from the analysis, and you should record how many were lost during conversion.

Understanding the Output

The result of enrichGO is an enrichResult object. The as.data.frame function converts it to a standard data frame for inspection and export:

ego_df <- as.data.frame(ego)

Each row represents one enriched GO term. The columns include the GO term ID, description, the gene ratio (proportion of input genes in the term), the background ratio (proportion of background genes in the term), the p-value, adjusted p-value, q-value, and the list of genes from your input that belong to the term.

The gene ratio and background ratio are the key numbers for interpretation. A term with a high gene ratio but low background ratio is specifically enriched in your gene list. A term with similar ratios in both is not enriched, even if it contains many of your genes.

Functional Class Scoring as an Alternative Approach

Over-representation analysis is the most common method, but it has a limitation: it requires a binary decision about which genes are significant. Genes that fall just below your threshold are excluded entirely. Functional class scoring (FCS) methods avoid this problem by using the full ranking of genes instead of a selected subset.

The bench scientist tutorial on functional class enrichment analysis describes both ORA and FCS as the two popular methods [<a href="#ref-1">1</a>]. FCS methods include gene set enrichment analysis (GSEA), which ranks all genes by a statistic such as log2 fold change or signal-to-noise ratio, then tests whether genes in a given functional category are concentrated at the top or bottom of the ranking.

clusterProfiler provides the gseGO function for GSEA-style analysis with GO terms:

gsea_go <- gseGO(
    geneList = ranked_genes,
    OrgDb = org.Hs.eg.db,
    keyType = "ENTREZID",
    ont = "BP",
    pAdjustMethod = "BH",
    pvalueCutoff = 0.05
)

The geneList argument requires a named numeric vector where the names are gene identifiers and the values are the ranking statistics. The vector must be sorted in decreasing order of the statistic.

The choice between ORA and FCS depends on your question. If you have a clear set of significant genes and want to know what functions they share, ORA is appropriate. If you have a continuous ranking and want to detect coordinated changes in pathways even when individual genes do not pass significance thresholds, FCS is more powerful. The tutorial on functional class enrichment analysis presents both methods with detailed R scripts for each task, including enrichment analysis, data processing, and visualization [<a href="#ref-1">1</a>].

Visualizing GO Enrichment Results

Dot Plot

The dot plot is the most common visualization for enrichment results. It displays GO terms on the y-axis, the gene ratio on the x-axis, and uses point size and color to show gene count and adjusted p-value. The dotplot function from enrichplot creates this figure:

dotplot(ego, showCategory = 20)

The showCategory argument controls how many terms to display. The terms are ordered by adjusted p-value, with the most significant at the top. This plot is useful for presentations and publications because it conveys the significance, effect size, and gene count for each term in a single figure.

Bar Plot

The bar plot is a simpler alternative that shows the count of genes or the enrichment score for each term:

barplot(ego, showCategory = 20)

This visualization is less information-dense than the dot plot but can be easier to read when the term names are long.

Term-Gene Network Plot

The term-gene network plot shows the connections between enriched GO terms and the genes that contribute to them. This plot is valuable for identifying genes that appear in multiple enriched terms, which may represent hub genes in the biological response:

cnetplot(ego, showCategory = 10)

The tutorial on functional class enrichment analysis presents annotated R code for multiple visualization types, including dot plots, term-gene network plots, enrichment map plots, ridge plots, and GSEA plots [<a href="#ref-1">1</a>]. Each visualization serves a different interpretive purpose, and you should generate several to understand your results fully.

Enrichment Map

The enrichment map plot shows the relationships between enriched terms based on the overlap of their gene sets. Terms that share many genes are connected by edges, revealing clusters of related biological functions:

emapplot(ego)

This plot is particularly useful when you have many enriched terms and need to identify the major themes in your data. The enrichment map can reveal that what appears to be many separate terms actually represents a few coordinated biological programs.

Interpreting Enrichment Results Correctly

Reading the Statistics

The p-value in enrichment analysis answers a specific question: if the background genes were randomly sampled, how often would we see this many genes from a particular GO term in a list of this size? A small p-value means the overlap between your gene list and the GO term is unlikely to be due to chance.

The adjusted p-value accounts for the fact that you are testing thousands of GO terms simultaneously. The Benjamini-Hochberg method controls the false discovery rate, which is the expected proportion of false positives among the terms you declare significant. The q-value is a similar measure that accounts for the fact that you are testing multiple ontologies and thresholds.

The gene ratio and background ratio provide effect size information. A term with a gene ratio of 0.2 and a background ratio of 0.02 means that 20 percent of your input genes belong to the term, while only 2 percent of the background genes do. This is a strong enrichment. A term with a gene ratio of 0.2 and a background ratio of 0.18 is not meaningfully enriched even if the p-value is small due to large sample sizes.

Handling Redundancy in GO Terms

The Gene Ontology is hierarchical, with parent terms that are more general and child terms that are more specific. This structure creates a redundancy problem: if a specific child term is enriched, its parent terms will also appear enriched because they contain the same genes. The result is a long list of overlapping terms that obscures the biological signal.

A recent method called evoGO addresses this problem by considering the GO hierarchy during the analysis [<a href="#ref-9">9</a>]. Built on conventional ORA, evoGO reduces the impact of differentially expressed genes on the significance of a given GO term if those genes already contribute to a higher significance of any descendant GO term [<a href="#ref-9">9</a>]. In synthetic benchmarks, evoGO reduced the number of enriched terms identified by ORA by an average of 30 percent while recovering the highest fraction of true positives [<a href="#ref-9">9</a>].

clusterProfiler provides a simpler approach to redundancy reduction through the simplify function, which removes terms that share a high proportion of genes with a more significant term:

ego_simplified <- simplify(ego, cutoff = 0.7, measure = "Wang")

The cutoff argument controls the similarity threshold. Higher values remove more terms. The measure argument specifies the semantic similarity measure used to compare terms.

Interpreting Results in Biological Context

Enrichment analysis identifies statistical associations, not causal mechanisms. An enriched term tells you that genes from your list are concentrated in a particular functional category, but it does not tell you whether that category is driving the phenotype you are studying.

The study of a CSF disease-associated macrophage signature in progressive multiple sclerosis illustrates how enrichment results are interpreted in context [<a href="#ref-10">10</a>]. The researchers identified CD14-positive monocytes in cerebrospinal fluid that resembled border-associated macrophages with a chronically activated antigen-presenting phenotype [<a href="#ref-10">10</a>]. They also found induction of disease-associated microglia and macrophage molecules, including TREM2, in secondary progressive MS [<a href="#ref-10">10</a>]. These enrichment findings supported the interpretation that specific cellular programs characterize different MS stages.

Similarly, a study on glioblastoma used enrichment analysis to validate a prognostic lncRNA model [<a href="#ref-11">11</a>]. The researchers identified differentially expressed genes, built a LASSO regression model with 19 lncRNAs, and validated the risk groups through enrichment analysis, immune profiling, and drug sensitivity evaluation [<a href="#ref-11">11</a>]. Functional enrichment highlighted lipid metabolism, JAK/STAT signaling, and T-cell regulation [<a href="#ref-11">11</a>]. These results connected the statistical model to biological mechanisms.

Working with Other Ontologies and Databases

KEGG Pathway Enrichment

GO terms describe gene functions, but many researchers also want to know which biological pathways are enriched. The Kyoto Encyclopedia of Genes and Genomes (KEGG) provides pathway annotations that complement GO. clusterProfiler includes the enrichKEGG function for this purpose:

ekegg <- enrichKEGG(
    gene = deg_entrez,
    organism = "hsa",
    pvalueCutoff = 0.05
)

The organism argument uses the KEGG organism code, such as "hsa" for human, "mmu" for mouse, or "ath" for Arabidopsis. The tutorial on functional class enrichment analysis notes that clusterProfiler can directly access the latest KEGG knowledge database [<a href="#ref-1">1</a>]. This direct access means you do not need to download and maintain separate KEGG annotation files.

The maize stalk study used KEGG enrichment alongside GO analysis to identify pathways related to hormone signal transduction, starch metabolism, and sucrose metabolism [<a href="#ref-3">3</a>]. The combination of GO and KEGG results provided complementary views of the biological response to planting density.

Disease Ontology Enrichment

Disease Ontology (DO) enrichment analysis connects your gene list to disease associations. The EnrichDO package, available through Bioconductor, implements a global weighted model that leverages the latest annotations of the human genome with DO terms and integrates DO graph topology on a global scale [<a href="#ref-12">12</a>]. Compared to classic enrichment methods, EnrichDO performs better in both GO and DO enrichment analysis cases and can accurately identify more specific terms without ignoring truly associated parent terms [<a href="#ref-12">12</a>].

The DOSE package in Bioconductor provides Disease Ontology enrichment within the same framework as clusterProfiler. The enrichDO function accepts the same gene vector and background arguments as enrichGO:

library(DOSE)
edo <- enrichDO(
    gene = deg_entrez,
    ont = "DO",
    pvalueCutoff = 0.05,
    readable = TRUE
)

Disease Ontology enrichment is particularly useful when you are studying a disease-related phenotype and want to confirm that your gene list is enriched for genes known to be associated with that disease. The EnrichDO publication demonstrates this with an Alzheimer's disease case where AD ranked first in the enrichment results [<a href="#ref-12">12</a>].

Reactome Pathway Enrichment

Reactome is a curated database of human biological pathways. The ReactomePA package provides enrichment analysis against this database:

BiocManager::install("ReactomePA")
library(ReactomePA)
ereact <- enrichPathway(
    gene = deg_entrez,
    organism = "human",
    pvalueCutoff = 0.05
)

The tutorial on functional class enrichment analysis lists GO, KEGG, and Reactome as the three basic knowledge databases used in the described methods [<a href="#ref-1">1</a>]. Each database provides a different perspective on gene function, and using all three can give a more complete picture than any single one.

Quality Control and Reproducibility

Recording Parameters and Versions

Enrichment results depend on many parameters: the annotation package version, the GO database version, the background gene set, the significance thresholds, and the multiple testing correction method. Any change in these parameters can change your results. For reproducible research, record all of them.

The Bioconductor project emphasizes reproducible genomic analysis in its official documentation [<a href="#ref-5">5</a>]. A reproducible workflow includes version information for R and all packages. The sessionInfo() function provides this information:

sessionInfo()

Include the output in your analysis report or supplementary materials. The nf-core documentation describes community standards for reproducible workflows, including version pinning and containerization [<a href="#ref-13">13</a>]. While nf-core is designed for pipeline development, the principles of version control and environment reproducibility apply to any analysis.

Validating Results with Independent Methods

A single enrichment analysis can produce false positives, particularly when the input gene list is large or the background set is poorly defined. Validate your key findings with independent methods. For example, if GO enrichment identifies a particular biological process, check whether the individual genes in that process show consistent expression changes in your data.

The spatial transcriptomics study using the CODA framework demonstrates this validation approach [<a href="#ref-14">14</a>]. After identifying spatially informative genes through enrichment analysis, the researchers used dual-color immunofluorescence experiments to confirm the findings [<a href="#ref-14">14</a>]. This experimental validation distinguishes robust enrichment signals from statistical artifacts.

Common Failure Patterns

Several recurring problems cause enrichment analysis to fail or produce misleading results. The most common is using an incorrect background set. If you use all genes in the genome as background when your experiment only measured a subset, you will overestimate significance for categories that are well represented on your platform. Always use the set of genes that were actually tested.

A second common failure is identifier mismatch. If your gene list uses symbols but the annotation package expects Entrez IDs, the analysis will fail or produce empty results. Always verify that your identifiers match the keyType argument.

A third failure pattern is interpreting enrichment results without considering redundancy. The hierarchical structure of GO means that many enriched terms are not independent. The evoGO method was developed specifically because redundancy in GO terms complicates prioritization and makes drawing concise conclusions challenging [<a href="#ref-9">9</a>]. Apply redundancy reduction before interpreting your results.

A fourth failure is ignoring the difference between statistical and biological significance. A term can be statistically enriched with a very small p-value but contain only a few genes with modest expression changes. Consider the effect size, represented by the gene ratio, alongside the p-value.

Limitations of GO Enrichment Analysis

Annotation Completeness

GO enrichment depends entirely on the quality and completeness of gene annotations. Genes that are poorly annotated or recently discovered may not appear in any GO terms, even if they are biologically important. This limitation is particularly relevant for non-model organisms. The FunFEA package for fungal genomics was developed because existing tools struggle with processing functional annotations of novel species for which no publicly available annotations exist [<a href="#ref-4">4</a>].

The NCBI databases are continuously updated as new annotations become available [<a href="#ref-7">7</a>]. Using the latest annotation packages ensures that your analysis reflects current knowledge, but it also means that results may change when annotations are updated. Record the annotation package version to ensure reproducibility.

Statistical Assumptions

Over-representation analysis assumes that genes are independent, but this assumption is violated in practice. Genes in the same pathway or chromosomal region are often co-regulated and may be correlated. This correlation can inflate significance for terms containing many correlated genes.

The choice of multiple testing correction also affects results. The Benjamini-Hochberg method controls the false discovery rate, which is appropriate for exploratory analysis. More conservative methods, such as Bonferroni correction, control the family-wise error rate but may miss true enrichments.

Interpretation Limits

Enrichment analysis identifies functional categories that are over-represented in your gene list, but it does not identify which genes are most important or which pathways are causal. A term can be enriched because of a single highly significant gene or because of many moderately significant genes. The term-gene network plot can help distinguish these scenarios by showing which genes contribute to each term.

The multiple sclerosis study illustrates the interpretive gap between enrichment and mechanism [<a href="#ref-10">10</a>]. The researchers identified a CSF disease-associated macrophage signature in progressive MS, but the enrichment results alone did not establish whether these macrophages drive progression or are a consequence of it [<a href="#ref-10">10</a>]. Additional experiments were needed to address causality.

Reporting Enrichment Results

Tables for Publication

The standard format for reporting enrichment results is a table with columns for the GO term ID, description, gene ratio, background ratio, p-value, adjusted p-value, and the contributing genes. Export the results to a CSV file for supplementary materials:

write.csv(ego_df, "go_enrichment_results.csv", row.names = FALSE)

The table should include all terms that passed your significance thresholds, beyond the top few. Reviewers may want to see the complete results to assess the specificity of the enrichment.

Figures for Presentations

The dot plot is the most effective single figure for communicating enrichment results. It shows the significance, effect size, and gene count for each term in a compact format. For presentations, limit the plot to the top 10 to 20 terms to keep the figure readable.

The term-gene network plot is useful for showing the relationship between genes and terms, but it becomes unreadable when too many terms are included. Limit the network plot to the top 5 to 10 terms.

Integrating with Other Analyses

Enrichment results are most powerful when integrated with other analyses. The glioblastoma study combined enrichment analysis with immune infiltration analysis, drug sensitivity evaluation, and survival analysis to build a complete picture of the prognostic model [<a href="#ref-11">11</a>]. The enrichment results identified the biological processes, while the other analyses connected those processes to clinical outcomes.

The maize stalk study integrated transcriptomics with targeted metabolomics [<a href="#ref-3">3</a>]. The transcriptome analysis identified differentially expressed genes enriched in pathways such as hormone signal transduction, starch metabolism, and sucrose metabolism [<a href="#ref-3">3</a>]. The metabolomics data confirmed that glucose was the primary sugar metabolite and showed how sucrose content changed with planting density [<a href="#ref-3">3</a>]. This integration of molecular data with functional enrichment provided a more complete understanding than either approach alone.

Practical Workflow Summary

Step 1: Prepare Your Gene List

Extract the significant genes from your differential expression results. Record your significance thresholds and the total number of genes tested. Convert gene identifiers to the type expected by your annotation package.

Step 2: Define Your Background Set

Use all genes that were tested for differential expression, beyond those that passed significance thresholds. For RNA-seq, this is typically all genes with sufficient expression. For microarrays, it is all genes on the array.

Step 3: Run Enrichment Analysis

Use enrichGO for GO terms, enrichKEGG for KEGG pathways, and enrichDO for Disease Ontology. Run each ontology separately to avoid mixing results. Record all parameters.

Step 4: Apply Redundancy Reduction

Use the simplify function or a method like evoGO to reduce redundant terms. Record the similarity threshold and measure used.

Step 5: Visualize Results

Generate dot plots, bar plots, term-gene network plots, and enrichment maps. Select the visualization that best communicates your findings.

Step 6: Interpret in Biological Context

Read the enriched terms as a group instead of individually. Identify the major themes and connect them to your experimental question. Validate key findings with independent methods.

Step 7: Report Reproducibly

Export results tables, save figures, and record session information. Include all parameters in your methods section.

At a Glance

Analysis StepKey FunctionCritical ParameterCommon Error
Identifier conversionbitrfromType and toType must match your data and annotation packageUsing symbols when the package expects Entrez IDs
GO enrichmentenrichGOuniverse must be all tested genes, not the whole genomeUsing the genome as background when only a subset was measured
KEGG enrichmentenrichKEGGorganism code must match your speciesUsing the wrong organism code produces empty results
Redundancy reductionsimplifycutoff controls how many terms are removedSkipping this step produces long lists of overlapping terms
Visualizationdotplot, cnetplot, emapplotshowCategory controls the number of terms displayedIncluding too many terms makes figures unreadable
ReproducibilitysessionInfoRecord all package versionsChanging annotation package versions changes results

Common Failure Patterns and Solutions

Empty Results

If enrichGO returns no enriched terms, check your input gene list and background set. The most common cause is an identifier mismatch. Verify that your gene identifiers match the keyType argument and that the annotation package contains annotations for your organism. Another cause is a background set that is too large or too small. The background should be the set of genes actually tested in your experiment.

Too Many Results

If you obtain hundreds of enriched terms, the problem is usually redundancy. The hierarchical structure of GO means that parent terms are enriched whenever child terms are enriched. Apply the simplify function or use a method like evoGO to reduce redundancy [<a href="#ref-9">9</a>]. Also consider whether your significance thresholds are too lenient. A q-value cutoff of 0.2 is common for exploratory analysis, but a more stringent cutoff may be appropriate for confirmatory work.

Results That Do Not Match Your Hypothesis

Enrichment analysis sometimes identifies terms that seem unrelated to your experimental question. This can happen when your gene list contains genes with multiple functions or when the background set is poorly defined. Before concluding that the analysis is wrong, examine the individual genes contributing to the unexpected terms. The term-gene network plot can help identify whether a few genes are driving the enrichment.

Results That Change Between Runs

If you rerun the analysis and get different results, the cause is usually a change in annotation packages or databases. The GO database is updated regularly, and new annotations can change enrichment results. The NCBI databases are continuously updated [<a href="#ref-7">7</a>]. Record the version of your annotation package and the date of your analysis to ensure reproducibility.

Professional Escalation Criteria

Some enrichment analysis problems require consultation with a bioinformatics specialist or statistician. Seek professional help in the following situations:

  • Your results change substantially when you update annotation packages, and you cannot determine which version is correct.
  • You are working with a non-model organism and need to create custom annotations.
  • Your enrichment results conflict with established biological knowledge, and you cannot resolve the discrepancy.
  • You need to integrate enrichment results across multiple omics data types and require statistical guidance.
  • Your study is intended to support regulatory decisions or clinical applications, and the analysis will be subject to external review.

The EMBL-EBI Training program provides learning pathways for bioinformatics analysis that can help you build the skills to address these challenges independently [<a href="#ref-15">15</a>]. The Galaxy Training Network offers accessible workflow training and analysis tutorials that complement the R-based approach described here [<a href="#ref-16">16</a>]. The Carpentries lessons provide foundational computing and data skills that support reproducible analysis [<a href="#ref-8">8</a>].

Frequently Asked Questions

What is the difference between over-representation analysis and functional class scoring?

Over-representation analysis (ORA) tests whether a selected list of significant genes contains more genes from a functional category than expected by chance. Functional class scoring (FCS) uses the full ranking of all genes and tests whether genes in a category are concentrated at the top or bottom of the ranking. ORA requires a binary decision about significance, while FCS uses the continuous ranking. The bench scientist tutorial on functional class enrichment analysis describes both methods as the two popular approaches [<a href="#ref-1">1</a>]. ORA is simpler and appropriate when you have a clear set of significant genes. FCS is more powerful for detecting coordinated changes in pathways where individual genes may not pass significance thresholds.

Why is the background gene set important in GO enrichment analysis?

The background set defines the expected frequency of each GO term. If you use all genes in the genome as background when your experiment only measured a subset, you will overestimate enrichment for categories that are well represented on your assay platform. The background should be all genes that were tested for differential expression. For RNA-seq, this is typically all genes with sufficient expression. For microarrays, it is all genes on the array. Using the correct background ensures that the enrichment statistics reflect your actual experimental design.

How do I choose between biological process, molecular function, and cellular component ontologies?

The choice depends on your research question. Biological process (BP) terms describe the larger programs that genes participate in, such as immune response or lipid metabolism. Molecular function (MF) terms describe the biochemical activities of gene products, such as kinase activity or DNA binding. Cellular component (CC) terms describe where gene products are located, such as nucleus or plasma membrane. Most researchers run BP as the primary analysis and use MF and CC for additional context. You can run all three simultaneously with ont = "ALL", but the results will be more complex to interpret.

What does the gene ratio mean in enrichment results?

The gene ratio is the proportion of your input genes that belong to a particular GO term. It is calculated as the number of input genes in the term divided by the total number of input genes. A high gene ratio means that a large fraction of your significant genes are involved in that function. The background ratio is the proportion of background genes in the term. The enrichment is the ratio of these two values. A term with a gene ratio of 0.2 and a background ratio of 0.02 represents a 10-fold enrichment.

How do I handle the redundancy of GO terms in my results?

The Gene Ontology is hierarchical, so specific child terms and their general parent terms often appear together in enrichment results. The simplify function in clusterProfiler removes terms that share a high proportion of genes with a more significant term. The evoGO method takes a different approach by considering the GO hierarchy during the analysis and reducing the impact of genes that already contribute to a more significant descendant term [<a href="#ref-9">9</a>]. In benchmarks, evoGO reduced the number of enriched terms by an average of 30 percent while recovering the highest fraction of true positives [<a href="#ref-9">9</a>].

Can I use clusterProfiler for non-model organisms?

Yes, but you may need to provide your own annotation. clusterProfiler relies on organism-level annotation packages that are available for common model organisms. For other species, you can create a custom annotation package or use a package like FunFEA that processes eggNOG-mapper annotations for novel genomes [<a href="#ref-4">4</a>]. FunFEA includes precomputed background frequency models for 65 commonly used fungal strains and supports GO, KEGG, and COG/KOG annotations [<a href="#ref-4">4</a>].

What is the difference between p-value, adjusted p-value, and q-value?

The p-value is the probability of observing the enrichment by chance, assuming the null hypothesis is true. The adjusted p-value accounts for multiple testing using a method such as Benjamini-Hochberg, which controls the false discovery rate. The q-value is an alternative measure that also accounts for multiple testing but uses a different estimation approach. In practice, the adjusted p-value and q-value are often similar. You should report the adjusted p-value in your methods and use it for significance decisions.

How should I report GO enrichment results in my publication?

Report the analysis parameters in your methods section, including the annotation package version, the background gene set, the significance thresholds, and the multiple testing correction method. Include a table of enriched terms in your supplementary materials with columns for the GO term ID, description, gene ratio, background ratio, adjusted p-value, and contributing genes. Include a dot plot or bar plot of the top terms in the main text. Record the session information to ensure reproducibility.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

[1] [Biological Functional Class Enrichment Analysis with R, an Annotated Tutorial for Bench Scientists.](https://pubmed.ncbi.nlm.nih.gov/41745221). Methods and protocols, 2026. [2] [Gene ontology enrichment analysis of congenital diaphragmatic hernia-associated genes](https://doi.org/10.1038/s41390-018-0192-8). Pediatric Research, 2018. [3] [Transcriptome and targeted metabolome analysis reveals sugar metabolism regulation in maize stalks at different densities.](https://doi.org/10.3389/fpls.2026.1784881). 2026. [4] [FunFEA: an R package for fungal functional enrichment analysis](https://doi.org/10.1186/s12859-025-06164-7). BMC Bioinformatics, 2025. [5] [Bioconductor](https://bioconductor.org/). Bioconductor Project. [6] [Using clusterProfiler to characterize multiomics data](https://doi.org/10.1038/s41596-024-01020-z). Nature Protocols, 2024. [7] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information. [8] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries. [9] [Cutting through the clutter: minimizing redundancy in GO enrichment analysis with evoGO](https://doi.org/10.1101/2025.02.24.639258). bioRxiv, 2025. [10] [A CSF disease-associated macrophage signature defines progressive multiple sclerosis.](https://doi.org/10.1186/s12974-026-03861-9). 2026. [11] [LncRNA SOX21-AS1 is associated with poor prognosis and immunomodulation in glioblastoma.](https://doi.org/10.21037/tcr-2025-1-2684). 2026. [12] [EnrichDO: a global weighted model for Disease Ontology enrichment analysis](https://doi.org/10.1093/gigascience/giaf021). GigaScience, 2025. [13] [nf-core Documentation](https://nf-co.re/docs). nf-core. [14] [Integrative cross-sample alignment and spatially differential gene analysis for spatial transcriptomics.](https://doi.org/10.1038/s41467-026-72862-2). 2026. [15] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute. [16] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project.

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.