Volcano Plots in RNA-seq: How to Visualize Differential Expression Results Effectively
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Volcano plots visualize differential gene expression by plotting log2 fold change (magnitude of change) against the negative log of the adjusted p-value (statistical significance), with genes showing large fold changes and high significance appearing in the upper corners.
- The y-axis must use adjusted p-values derived from multiple testing correction (e.g., Benjamini-Hochberg) to accurately reflect statistical confidence and avoid overstating significance, unlike raw p-values.
- Thresholds for log2 fold change (typically 1 or 2) and adjusted p-value (typically 0.05) should be biologically relevant and clearly stated, as they define which genes are highlighted as candidates for further investigation.
- Selecting genes based on both statistical significance and fold change (double filtering) can inflate the false discovery rate; genes identified should be treated as candidates requiring validation via independent methods like RT-qPCR.
- Transparently documenting analysis parameters, including software versions, thresholds, labeling strategies, and axis transformations, is crucial for ensuring the reproducibility of volcano plots.
- Interpreting volcano plots requires integrating findings with functional enrichment analyses (e.g., Gene Ontology) and validating candidate genes with independent experimental methods to connect visualization to biological meaning and disease mechanisms.
Volcano plots are a standard visualization method in RNA-seq differential expression analysis that display statistical significance against magnitude of change for every detected gene. A properly constructed volcano plot places genes with large fold changes and small adjusted p-values in the upper corners of the plot, making candidate differentially expressed genes immediately visible. This article explains how to build publication-ready volcano plots using EnhancedVolcano in R, how to label top genes, how to set meaningful thresholds, and how to avoid the common visualization and statistical mistakes that lead to misleading figures.
The intended reader is a biology student, researcher, or laboratory professional who has completed differential expression analysis and needs to present results clearly. The practical outcome is a reproducible workflow for generating volcano plots that accurately represent your data and support your biological conclusions.
What a Volcano Plot Shows in Differential Expression Analysis
A volcano plot is a scatter plot with two axes. The x-axis represents the log2 fold change between two experimental conditions, such as treated versus control or disease versus healthy. The y-axis represents the negative logarithm of the adjusted p-value, which means genes with smaller p-values appear higher on the plot. Each point on the plot corresponds to one gene or transcript.
The name comes from the visual shape. Genes with large fold changes and high statistical significance form two distinct arms that rise from the central cloud of points, resembling a volcanic eruption. Genes that do not change between conditions cluster near zero on the x-axis and low on the y-axis.
The plot serves a practical filtering function. After applying a multiple testing correction such as the Benjamini-Hochberg procedure, researchers often still have thousands of significant genes. The volcano plot helps prioritize which genes merit closer inspection by combining two criteria: effect size and statistical significance. Genes in the upper left corner are significantly downregulated, and genes in the upper right corner are significantly upregulated.
Published studies across many fields use volcano plots to present differential expression results. For example, a study of long non-coding RNA in psoriasis used a volcano plot to display 412 upregulated and 625 downregulated differentially expressed lncRNAs in psoriatic tissue compared with normal skin tissue [<a href="#ref-1">1</a>]. A study of histone deacetylase 4 in chondrocytes used volcano plot analysis to identify differentially expressed genes, finding 1,483 upregulated and 1,185 downregulated genes with a p-value threshold below 0.05 [<a href="#ref-2">2</a>]. These examples show that volcano plots are a routine part of the RNA-seq reporting workflow.
At a Glance
| Decision Point | Recommended Practice | Common Mistake |
|---|---|---|
| Y-axis values | Use adjusted p-values after multiple testing correction | Plotting raw p-values overstates significance |
| Fold change threshold | Set based on biological relevance, typically log2 fold change of 1 or 2 | Using zero threshold labels genes with trivial effect sizes |
| Gene labeling | Limit to top 10 to 20 genes or a custom biological list | Labeling all significant genes creates unreadable overlap |
| Statistical interpretation | Treat volcano plot selections as candidates for validation | Assuming double-filtered genes are confirmed discoveries |
| Axis scaling | Cap extreme values transparently and state limits in legend | Hiding axis caps misleads readers about data range |
| Tool selection | Use EnhancedVolcano in R, ExpressAnalyst, or TCC-GUI based on skill level | Choosing a tool without considering reproducibility needs |
Core Principles of Volcano Plot Construction
The Relationship Between Fold Change and Statistical Significance
The two dimensions of a volcano plot encode different information. The log2 fold change measures the magnitude of expression difference between conditions. A log2 fold change of 1 means a doubling of expression, and a log2 fold change of negative 1 means a halving of expression. The adjusted p-value measures the statistical confidence that the observed difference is not due to random variation.
These two measures are independent. A gene can have a large fold change but fail to reach statistical significance because of high variability between replicates. Another gene can have a small fold change but achieve high significance because expression is highly consistent across replicates. The volcano plot makes this distinction visible, which is why it is more informative than a simple list of significant genes.
Threshold Selection and Its Consequences
Thresholds determine which genes are highlighted as differentially expressed. The two common thresholds are the adjusted p-value cutoff and the log2 fold change cutoff. A typical adjusted p-value threshold is 0.05, and a typical log2 fold change threshold is 1 or 2, corresponding to a 2-fold or 4-fold change in expression.
Threshold selection should reflect your experimental context. A discovery-oriented study with no prior hypothesis might use a more permissive fold change threshold to avoid missing candidate genes. A validation study testing a specific mechanism might use a stricter threshold to reduce false positives. The thresholds you choose should be stated clearly in the figure legend and in the methods section of any report.
The Role of Multiple Testing Correction
RNA-seq experiments measure expression for tens of thousands of genes simultaneously. Testing each gene for differential expression creates a multiple testing problem. Without correction, thousands of genes would appear significant by chance alone. The Benjamini-Hochberg procedure controls the false discovery rate, which is the expected proportion of false positives among the genes declared significant.
Volcano plots typically use the adjusted p-value instead of the raw p-value on the y-axis. This practice ensures that the genes highlighted as significant have passed multiple testing correction. Using raw p-values on a volcano plot overstates significance and can lead to incorrect biological conclusions.
Statistical Limitations of Volcano Plot Filtering
The Double Filtering Problem
A volcano plot visually suggests a double filtering procedure: select genes that pass both the adjusted p-value threshold and the fold change threshold. This approach is intuitive, but it has a statistical flaw. The Benjamini-Hochberg procedure guarantees false discovery rate control over the full set of discoveries, but it does not guarantee error control over a filtered subset.
Research published in Briefings in Bioinformatics demonstrated this problem directly. The authors showed that selecting features with both small adjusted p-values and large estimated effect sizes can produce an inflated number of false discoveries [<a href="#ref-3">3</a>]. Their simulation experiments and RNA-seq data analysis revealed that the feature with the largest estimated effect is a very likely false positive result [<a href="#ref-3">3</a>].
This finding has practical implications. If you use a volcano plot to select the top genes for follow-up validation, you may be selecting genes that are statistical artifacts. The genes with the most extreme fold changes are not necessarily the most reliable discoveries.
Alternative Approaches to Double Filtering
The same research group investigated alternative multiple testing procedures that do not inflate the false discovery rate when double filtering is applied [<a href="#ref-3">3</a>]. These procedures account for the fact that filtering on effect size after multiple testing correction changes the error rate. The authors implemented their procedure in an interactive web application that is publicly available [<a href="#ref-3">3</a>].
For most researchers, the practical response is to be cautious when interpreting volcano plot selections. The volcano plot is a visualization tool, not a statistical inference procedure. Genes selected from a volcano plot should be treated as candidates for validation, not as confirmed discoveries. Validation can take the form of RT-qPCR, as demonstrated in studies that confirmed RNA-seq findings with independent measurements [<a href="#ref-1">1</a>][<a href="#ref-2">2</a>].
Preparing Your Data for Volcano Plot Generation
Required Input Data Structure
EnhancedVolcano in R requires a data frame with specific columns. The minimal requirement is a column for gene identifiers, a column for log2 fold change, and a column for adjusted p-values. The gene identifier column is typically named something like "gene" or "symbol". The log2 fold change column is often named "log2FoldChange" or "logFC". The adjusted p-value column is often named "padj" or "adj.P.Val".
The data frame should be the output of a differential expression analysis tool such as DESeq2, edgeR, or limma. These tools produce results tables that contain the necessary columns. You may need to convert the results table to a data frame and ensure that gene identifiers are present as a column instead of as row names.
Data Quality Checks Before Plotting
Before generating a volcano plot, verify that your input data meets basic quality standards. Check that the log2 fold change column contains numeric values and that no infinite values are present. Check that the adjusted p-value column contains values between 0 and 1. Check that gene identifiers are unique and that no duplicate identifiers exist.
Missing values require attention. Some differential expression tools report NA for adjusted p-values when a gene is filtered out during analysis. Decide how to handle these values before plotting. Options include removing genes with missing adjusted p-values or keeping them in the plot but not labeling them as significant.
Handling Genes with Extreme Values
RNA-seq data often contains genes with very large fold changes or very small p-values. These extreme values can compress the visual scale of the volcano plot, making it difficult to see the distribution of the majority of genes. EnhancedVolcano provides options for handling this issue, such as capping the axes or using a transformed scale.
The decision to cap axes should be made transparently. If you limit the x-axis to a maximum absolute log2 fold change, state this in the figure legend. If you limit the y-axis to a maximum negative log p-value, state this as well. Capping axes changes the visual appearance of the plot but does not change the underlying data.
Using EnhancedVolcano in R
Installation and Basic Usage
EnhancedVolcano is an R package available through Bioconductor. Bioconductor provides official documentation for package installation and usage, and it is the standard repository for genomic analysis software in R [<a href="#ref-4">4</a>]. The package is designed to create publication-ready volcano plots with minimal code.
The basic function call takes the data frame, the gene identifier column name, the log2 fold change column name, and the adjusted p-value column name as arguments. Additional arguments control labeling, thresholds, and aesthetics. The function returns a ggplot object that can be customized further or saved directly to a file.
Labeling Top Genes
One of the main advantages of EnhancedVolcano over base R plotting is the ability to label genes directly on the plot. The labeling can be controlled in several ways. You can label all genes that pass both thresholds, which may produce a crowded plot if many genes are significant. You can label a specific number of top genes based on adjusted p-value or fold change. You can also provide a custom list of gene names to label.
The choice of labeling strategy depends on the purpose of the plot. A plot intended for a publication figure might label only the top 10 or 20 genes to keep the figure readable. A plot intended for exploratory analysis might label all significant genes so that the researcher can identify candidates visually.
Adjusting Significance Thresholds
EnhancedVolcano allows you to set the adjusted p-value cutoff and the log2 fold change cutoff independently. The default adjusted p-value cutoff is typically 0.05, and the default log2 fold change cutoff is typically 1. These defaults can be changed to match your analysis parameters.
When you change the thresholds, the plot updates the coloring of points. Genes passing both thresholds are colored distinctly, typically in red. Genes passing only the p-value threshold or only the fold change threshold are colored in different shades. Genes passing neither threshold remain in the background color.
Customizing Plot Aesthetics
Publication-ready figures require attention to aesthetics. EnhancedVolcano provides arguments for controlling point size, point shape, point transparency, and color scheme. You can also control the axis labels, the plot title, and the legend position.
The default color scheme uses red for significant genes, blue for genes passing only the p-value threshold, green for genes passing only the fold change threshold, and grey for non-significant genes. These colors can be customized to match journal requirements or to accommodate color-blind readers.
Practical Workflow for Generating a Volcano Plot
Step 1: Load Your Differential Expression Results
Start by loading your differential expression results into R. The results should be in a data frame with columns for gene identifiers, log2 fold change, and adjusted p-value. If your results are in a CSV or TSV file, use the appropriate read function to import them.
Verify the structure of your data frame before proceeding. Use functions to check the column names, the number of rows, and the presence of missing values. This verification step prevents errors later in the workflow.
Step 2: Install and Load EnhancedVolcano
Install EnhancedVolcano from Bioconductor if it is not already installed. The installation command is documented on the Bioconductor website [<a href="#ref-4">4</a>]. After installation, load the package with the library function.
If you are working in an environment where you cannot install packages, consider using an alternative tool. ExpressAnalyst is a web-based platform that can generate volcano plots without requiring local R installation [<a href="#ref-5">5</a>]. The platform supports end-to-end transcriptomics analysis, including differential expression and visualization [<a href="#ref-5">5</a>].
Step 3: Generate the Basic Plot
Create a basic volcano plot with the EnhancedVolcano function. Provide the data frame and the column names as arguments. Examine the resulting plot to ensure that the axes are oriented correctly and that the points are distributed as expected.
The basic plot should show the characteristic volcano shape. If the plot looks unusual, check your data for errors. Common problems include incorrect column names, log2 fold changes that are actually log10 fold changes, and adjusted p-values that are actually raw p-values.
Step 4: Adjust Thresholds and Labeling
Set the adjusted p-value cutoff and the log2 fold change cutoff to match your analysis parameters. Choose a labeling strategy that suits the purpose of your plot. For a publication figure, label a limited number of top genes. For exploratory analysis, label all significant genes or provide a custom gene list.
Review the labeled genes to ensure that they are biologically plausible. If a labeled gene is clearly an artifact, such as a ribosomal RNA or a known technical artifact, consider whether the labeling strategy needs adjustment.
Step 5: Customize Aesthetics and Export
Customize the plot aesthetics to meet your requirements. Adjust point size, transparency, colors, axis labels, and title. Ensure that the plot is legible at the size it will be reproduced in your target venue.
Export the plot in a suitable format. For publication, a vector format such as PDF or SVG is preferred because it scales without loss of quality. For web use, a high-resolution PNG or TIFF is acceptable. Set the resolution to at least 300 DPI for print reproduction.
Alternative Tools for Volcano Plot Generation
Web-Based Platforms
Researchers who do not use R can generate volcano plots through web-based platforms. ExpressAnalyst is a web-based platform that enables intuitive end-to-end transcriptomics and proteomics data analysis [<a href="#ref-5">5</a>]. Users can start from FASTQ files, gene or protein abundance tables, or gene or protein lists [<a href="#ref-5">5</a>]. The platform performs read quantification, expression table processing and normalization, and differential expression analysis [<a href="#ref-5">5</a>]. Results are presented through interactive visualizations including volcano plots, heatmaps, networks, and ridgeline charts [<a href="#ref-5">5</a>].
ExpressAnalyst supports 29 common organisms and can perform transcriptome profiling for non-model organisms without good reference genomes [<a href="#ref-5">5</a>]. The platform is organized into 11 basic protocols covering tasks from count table uploading to dose-response analysis [<a href="#ref-5">5</a>]. This structure makes it suitable for researchers who need a guided workflow.
Shiny-Based Applications
TCC-GUI is a Shiny-based application for differential expression analysis of RNA-seq count data [<a href="#ref-6">6</a>]. The application provides a graphical user interface that requires only a web browser [<a href="#ref-6">6</a>]. It includes differential expression pipelines with robust normalization and simulation data generation [<a href="#ref-6">6</a>]. The visualization tools include volcano plots and heatmaps with hierarchical clustering [<a href="#ref-6">6</a>].
TCC-GUI is designed for non-R users who need access to differential expression analysis functionality [<a href="#ref-6">6</a>]. The application generates reports using R Markdown, which supports reproducible documentation of the analysis [<a href="#ref-6">6</a>]. The source code is available on GitHub under the MIT license [<a href="#ref-6">6</a>].
Choosing Between Tools
The choice of tool depends on your context. If you are already working in R and have differential expression results, EnhancedVolcano is a direct and flexible option. If you prefer a graphical interface or need to analyze data from non-model organisms, ExpressAnalyst provides a structured workflow [<a href="#ref-5">5</a>]. If you need a simple interface for differential expression analysis without R knowledge, TCC-GUI offers a browser-based solution [<a href="#ref-6">6</a>].
Regardless of the tool you choose, the underlying principles of volcano plot construction remain the same. You need accurate differential expression results, appropriate thresholds, and transparent reporting of your visualization choices.
Tool Comparison for Volcano Plot Generation
| Feature | EnhancedVolcano | ExpressAnalyst | TCC-GUI |
|---|---|---|---|
| Required skill level | R programming knowledge | Web browser navigation | Web browser navigation |
| Installation | Bioconductor package [<a href="#ref-4">4</a>] | None, web-based [<a href="#ref-5">5</a>] | None, web-based [<a href="#ref-6">6</a>] |
| Input format | Data frame with log2 fold change and adjusted p-value | FASTQ files, abundance tables, or gene lists [<a href="#ref-5">5</a>] | RNA-seq count data [<a href="#ref-6">6</a>] |
| Customization | High, full ggplot control | Interactive visualization options [<a href="#ref-5">5</a>] | GUI-based options [<a href="#ref-6">6</a>] |
| Reproducibility | Script-based, version controllable | Protocol-based workflow [<a href="#ref-5">5</a>] | R Markdown reporting [<a href="#ref-6">6</a>] |
| Best use case | Publication figures with custom labeling | End-to-end analysis for non-R users [<a href="#ref-5">5</a>] | Differential expression with normalization [<a href="#ref-6">6</a>] |
Common Failure Patterns in Volcano Plot Generation
Using Raw P-Values Instead of Adjusted P-Values
A frequent mistake is plotting raw p-values on the y-axis instead of adjusted p-values. Raw p-values do not account for multiple testing and will overstate the number of significant genes. The resulting volcano plot will show many more points in the significant regions than are actually supported by the data.
The fix is to use adjusted p-values in the plot. Most differential expression tools provide adjusted p-values in their output. If your tool does not, apply a multiple testing correction before plotting.
Setting Thresholds That Are Too Permissive or Too Strict
Threshold selection requires judgment. A log2 fold change threshold of 0 means that any gene with a significant adjusted p-value is highlighted, regardless of effect size. This approach produces a plot where many genes with trivial fold changes are labeled as significant. A log2 fold change threshold of 5 means that only genes with a 32-fold change are highlighted, which may exclude biologically important genes with smaller but real effects.
The appropriate threshold depends on your biological question. Consider the expected magnitude of expression changes in your system. Consider the variability between replicates. Consider the consequences of false positives versus false negatives for your downstream validation plans.
Ignoring the Double Filtering Problem
Selecting genes from a volcano plot by requiring both a small adjusted p-value and a large fold change can inflate the false discovery rate [<a href="#ref-3">3</a>]. This problem is not widely appreciated, but it has been demonstrated in simulation experiments and real RNA-seq data [<a href="#ref-3">3</a>].
The practical response is to treat volcano plot selections as candidates instead of confirmed discoveries. Validate selected genes with an independent method such as RT-qPCR. Report the validation results transparently, including any genes that do not confirm.
Overcrowding the Plot with Labels
Labeling every significant gene on a volcano plot produces an unreadable figure. When hundreds or thousands of genes pass the thresholds, the labels overlap and obscure the underlying data distribution. The plot becomes visually confusing and fails to communicate the key findings.
The fix is to limit labeling. Label only the top genes by adjusted p-value or by absolute fold change. Alternatively, label a custom list of genes that are biologically relevant to your study. The figure legend should state the labeling criterion.
Misinterpreting the Absence of Significance
A gene that does not pass the significance thresholds is not necessarily unimportant. The gene may have a real biological effect that is too small to detect with the available sample size. The gene may have high variability between replicates that masks a consistent difference. The volcano plot shows what the experiment can detect, not the complete biological reality.
Report the limitations of your experiment alongside the volcano plot. State the sample size, the variability, and the detection power. This context helps readers interpret the plot correctly.
Failing to Document Axis Transformations
When you cap axes or apply transformations to handle extreme values, you change how readers perceive the data distribution. A plot with a capped y-axis may make many genes appear equally significant when their actual p-values differ by orders of magnitude. Without documentation, readers cannot know that the visual scale is compressed.
Always state axis limits and transformations in the figure legend. If you use a log scale or a capped range, explain what the reader is seeing. This transparency is essential for accurate interpretation.
Records and Measurements for Reproducible Volcano Plots
Documenting Analysis Parameters
Reproducible volcano plots require documentation of all analysis parameters. Record the version of the differential expression tool, the version of EnhancedVolcano, and the version of R. Record the thresholds used for adjusted p-value and log2 fold change. Record the labeling strategy and the axis limits.
This documentation should be stored with the analysis code and the input data. A collaborator or reviewer should be able to reproduce your volcano plot exactly by following the documented parameters.
Recording Gene Counts at Each Threshold
A useful record is the number of genes passing each threshold combination. Count the genes that pass the adjusted p-value threshold alone, the fold change threshold alone, and both thresholds. Record these counts in your analysis notebook or methods section.
These counts provide context for the volcano plot. A plot with 5,000 significant genes tells a different story than a plot with 50 significant genes. The counts also help readers understand the stringency of your thresholds.
Maintaining Version Control
Version control is essential for reproducible analysis. Use a version control system such as Git to track changes to your analysis code. The Carpentries provides lessons on version control with Git that are suitable for researchers who need to learn these skills [<a href="#ref-7">7</a>].
Version control allows you to revisit previous versions of your analysis and understand how the volcano plot changed as the analysis evolved. This record is valuable for troubleshooting and for responding to reviewer questions.
Recording Validation Outcomes
When you validate volcano plot selections with RT-qPCR or another method, record the outcomes systematically. Note which genes confirmed the RNA-seq direction of change, which genes did not confirm, and which genes showed discordant results. This record is valuable for assessing the reliability of your differential expression analysis.
A study of long non-coding RNAs in psoriasis provides a model for this practice. The researchers selected six candidate lncRNAs from RNA-seq and volcano plot analysis, then validated all six by RT-qPCR in 25 pairs of samples [<a href="#ref-1">1</a>]. Three lncRNAs were confirmed as increased in psoriatic tissue, and three were confirmed as decreased [<a href="#ref-1">1</a>]. This transparent reporting of validation outcomes strengthens the study conclusions.
Quality Controls for Volcano Plot Interpretation
Checking the Distribution of Points
Before interpreting a volcano plot, examine the overall distribution of points. A healthy volcano plot shows a dense cloud of non-significant genes near zero fold change and low significance. The significant genes should form two arms extending upward and outward.
If the plot shows an unusual pattern, investigate the cause. A plot with no significant genes may indicate a weak biological effect or a problem with the analysis. A plot with too many significant genes may indicate a batch effect or an error in the experimental design.
Verifying Known Controls
If your experiment includes genes with known expression behavior, verify that these genes appear in the expected regions of the volcano plot. A housekeeping gene that should not change between conditions should appear near zero fold change. A gene that is known to be induced by your treatment should appear in the upregulated arm.
This verification step catches systematic errors in the analysis. If known controls appear in unexpected locations, the differential expression analysis may have a problem that needs correction before the volcano plot is interpreted.
Comparing Replicates for Consistency
The volcano plot is only as reliable as the underlying differential expression analysis. Before generating the plot, verify that biological replicates are consistent. Principal component analysis can reveal whether replicates cluster by condition, as demonstrated in a study that used PCA to show that lncRNA profiles clearly distinguished psoriatic tissue from normal tissue [<a href="#ref-1">1</a>].
If replicates do not cluster by condition, the differential expression results may be unreliable. Investigate the cause before proceeding to visualization. Possible causes include sample mislabeling, batch effects, and technical variation.
Assessing Power and Detection Limits
A volcano plot cannot show genes that were not detected in your experiment. Lowly expressed genes may be absent from the analysis entirely. Genes with high technical variability may fail to reach significance even when biological effects exist. Understanding your experiment detection limits helps you interpret the plot accurately.
Consider how your sequencing depth and sample size affect your ability to detect differential expression. A study with few replicates will have less power to detect small fold changes. The volcano plot will show fewer significant genes, but this reflects the experimental design, not the absence of biological differences.
Interpreting Volcano Plots in Biological Context
Moving from Visualization to Biological Meaning
A volcano plot identifies candidate genes, but it does not explain their biological significance. The next step is functional analysis. Gene Ontology enrichment analysis can reveal which biological processes are overrepresented among the differentially expressed genes. Pathway analysis can reveal which signaling pathways are affected.
Published studies routinely combine volcano plots with functional analysis. A study of rheumatoid arthritis remission used volcano plots to display differentially expressed genes and then performed Gene Ontology enrichment analysis to understand the biological functions of those genes [<a href="#ref-8">8</a>]. A study of porcine epidemic diarrhea virus infection identified differentially expressed genes and then used Gene Ontology and KEGG analysis to identify enriched pathways related to immune response and viral replication [<a href="#ref-9">9</a>].
Integrating Multiple Data Types
Volcano plots from bulk RNA-seq can be integrated with other data types to strengthen biological conclusions. Single-cell RNA-seq can reveal which cell types express the differentially expressed genes. Proteomics can confirm that changes in mRNA expression lead to changes in protein abundance. Spatial transcriptomics can reveal where in the tissue the expression changes occur.
A study of acute kidney injury integrated bulk RNA-seq, proteomic, and single-cell RNA-seq datasets to identify key transcription factors [<a href="#ref-10">10</a>]. The integration revealed that loss of KLF15 expression in proximal tubule cells was associated with a proinflammatory and pro-apoptotic state [<a href="#ref-10">10</a>]. This type of integration provides more confidence than a single data type alone.
Validating Candidate Genes
Genes selected from a volcano plot should be validated with an independent method. RT-qPCR is the standard validation approach for mRNA expression changes. A study of long non-coding RNAs in psoriasis used RT-qPCR to validate six candidate lncRNAs identified from RNA-seq and volcano plot analysis [<a href="#ref-1">1</a>]. Three lncRNAs were confirmed as increased in psoriatic tissue, and three were confirmed as decreased [<a href="#ref-1">1</a>].
Validation serves two purposes. It confirms that the RNA-seq measurements are accurate. It also provides an independent assessment of the biological reality of the expression changes. Genes that fail validation may be technical artifacts or may reflect differences between the RNA-seq and RT-qPCR measurement methods.
Connecting to Disease Mechanisms
Volcano plots can identify genes that connect to disease mechanisms when interpreted in biological context. A study of breast cancer used integrated analysis to identify GREB1 as a key effector regulated by a BRD4-associated super-enhancer [<a href="#ref-11">11</a>]. The study showed that a BRD4 PROTAC degrader enhanced fulvestrant sensitivity through GREB1 signaling disruption [<a href="#ref-11">11</a>]. This type of mechanistic interpretation moves beyond the volcano plot to generate testable hypotheses.
Similarly, a study of hepatocellular carcinoma identified SF3B4 as a splicing factor that stabilizes SREBF1 mRNA through 3'UTR binding [<a href="#ref-12">12</a>]. The study used transcriptomic analyses to show that SF3B4 knockdown induced widespread transcriptional remodeling [<a href="#ref-12">12</a>]. These examples illustrate how volcano plot candidates can lead to mechanistic insights when followed with functional experiments.
Safety and Regulatory Context for RNA-seq Visualization
Data Management and Privacy
RNA-seq data may contain sensitive information, particularly if it comes from human subjects. Researchers must follow institutional and legal requirements for data storage, sharing, and publication. The NCBI provides data resources and search systems for sequence data, and researchers should be aware of the data submission and access policies for these resources [<a href="#ref-13">13</a>].
When publishing volcano plots, consider whether the gene-level data reveals sensitive information. In most cases, aggregated gene expression data does not identify individual subjects. However, the raw sequence data may contain identifiable information and should be handled according to applicable regulations.
Reproducibility Requirements
Many journals and funding agencies require that analysis code and data be made available for reproducibility. The Galaxy Training Network provides accessible workflow training that emphasizes reproducibility in bioinformatics analysis [<a href="#ref-14">14</a>]. The nf-core documentation describes community pipeline standards that support reproducible workflow configuration [<a href="#ref-15">15</a>].
A reproducible volcano plot workflow includes the input data, the analysis code, the software versions, and the parameter settings. This package of materials allows others to verify the results and to apply the same methods to their own data.
Ethical Reporting of Results
Volcano plots can create an impression of certainty that the underlying data may not support. The double filtering problem can produce inflated false discovery rates [<a href="#ref-3">3</a>]. The largest effect sizes are particularly likely to be false positives [<a href="#ref-3">3</a>].
Report volcano plot results with appropriate caution. State the limitations of the analysis. Describe the validation steps that were or were not performed. Avoid overstating the significance of individual genes selected from the plot.
Training and Competency Considerations
Generating and interpreting volcano plots requires foundational bioinformatics skills. The EMBL-EBI provides training pathways for bioinformatics data resources and practical analysis education [<a href="#ref-16">16</a>]. The Carpentries offers lessons on foundational computing, data, shell, Git, and programming that support reproducible analysis practices [<a href="#ref-7">7</a>].
Researchers should ensure they have adequate training before performing differential expression analysis and visualization. Misinterpretation of volcano plots can lead to incorrect biological conclusions and wasted validation effort. Seeking training or collaboration with experienced bioinformaticians is appropriate when skills are lacking.
Professional Escalation Criteria
When to Seek Statistical Consultation
If your volcano plot shows an unexpected pattern that you cannot explain, seek statistical consultation. A biostatistician or bioinformatics specialist can review your analysis and identify potential problems. This consultation is particularly important if the results will be used for clinical decisions or regulatory submissions.
Signs that consultation is needed include an unusually high number of significant genes, a complete absence of significant genes, or significant genes that do not make biological sense. These patterns may indicate problems with the experimental design, the data processing, or the statistical analysis.
When to Revisit the Experimental Design
If the volcano plot reveals that your experiment lacks power to detect meaningful differences, the experimental design may need revision. Increasing the number of biological replicates is the most direct way to increase power. Reducing technical variability through improved protocols can also help.
The decision to repeat an experiment should be made in consultation with collaborators and, if applicable, with funding agencies. Repeating an experiment has costs in time and resources, but it may be necessary to obtain reliable results.
When to Escalate Data Quality Issues
If quality checks reveal problems with the input data, escalate the issue before proceeding with visualization. Problems such as sample mislabeling, batch effects, and contamination can invalidate the differential expression results. The volcano plot will faithfully display the results of a flawed analysis, which is worse than no plot at all.
The NCBI provides resources for understanding sequence data quality and for accessing reference data [<a href="#ref-13">13</a>]. The EMBL-EBI provides training materials for bioinformatics analysis that include guidance on data quality assessment [<a href="#ref-16">16</a>]. These resources can help you identify and address data quality issues.
When to Consult Domain Experts
If your volcano plot identifies candidate genes that are unexpected based on your biological knowledge, consult a domain expert before proceeding with extensive validation. The expert may recognize patterns that indicate technical artifacts or may suggest alternative interpretations of the data.
A study of intestinal epithelial cells provides an example of domain-specific interpretation. The researchers used bulk RNA sequencing to investigate how biomimetic substrates and shear stress affect intestinal stem cell-derived monolayers [<a href="#ref-17">17</a>]. Their analysis revealed that dynamic culture conditions activated metabolic and immune pathways [<a href="#ref-17">17</a>]. This type of interpretation benefits from expertise in intestinal biology and tissue engineering.
Frequently Asked Questions
What is the difference between raw p-value and adjusted p-value on a volcano plot?
The raw p-value is the probability of observing the data if there is no true difference in expression. The adjusted p-value accounts for the fact that thousands of genes are tested simultaneously. Volcano plots should use adjusted p-values because raw p-values overstate significance. Using raw p-values will show many more genes as significant than are actually supported by the data.
How do I choose the log2 fold change threshold for my volcano plot?
The log2 fold change threshold should reflect the magnitude of change that is biologically meaningful in your system. A threshold of 1 corresponds to a 2-fold change, and a threshold of 2 corresponds to a 4-fold change. Consider the expected effect sizes in your experiment and the consequences of missing real changes versus including false positives. State your chosen threshold in the figure legend.
Why does my volcano plot show no significant genes?
A volcano plot with no significant genes may indicate a weak biological effect, high variability between replicates, or insufficient sample size. Check that the differential expression analysis was performed correctly and that the adjusted p-values are being used. Consider whether the experimental design has adequate power to detect the expected effect sizes.
How many genes should I label on a volcano plot?
The number of labeled genes depends on the purpose of the plot. For a publication figure, labeling the top 10 to 20 genes by adjusted p-value or absolute fold change keeps the figure readable. For exploratory analysis, you may label more genes or provide a custom list of biologically relevant genes. Overcrowding the plot with labels makes it unreadable.
Can I use a volcano plot to select genes for validation?
You can use a volcano plot to identify candidate genes for validation, but you should be aware of the statistical limitations. Selecting genes with both small adjusted p-values and large fold changes can inflate the false discovery rate [<a href="#ref-3">3</a>]. Treat volcano plot selections as candidates and validate them with an independent method such as RT-qPCR.
What is the double filtering problem in volcano plots?
The double filtering problem occurs when you select genes that pass both an adjusted p-value threshold and a fold change threshold. The Benjamini-Hochberg procedure controls the false discovery rate over the full set of discoveries, but it does not guarantee error control over a filtered subset [<a href="#ref-3">3</a>]. The selected subset may contain an inflated number of false discoveries [<a href="#ref-3">3</a>].
What tools can I use if I do not know R?
If you do not use R, you can generate volcano plots through web-based platforms. ExpressAnalyst is a web-based platform that supports end-to-end transcriptomics analysis and generates volcano plots as part of its visualization options [<a href="#ref-5">5</a>]. TCC-GUI is a Shiny-based application that provides a graphical interface for differential expression analysis and includes volcano plot visualization [<a href="#ref-6">6</a>].
How should I report volcano plot parameters in my methods section?
Report the software and version used to generate the plot, the adjusted p-value threshold, the log2 fold change threshold, the labeling strategy, and any axis limits. Report the number of genes passing each threshold. This information allows readers to understand exactly what the plot shows and to reproduce it if needed.
Related Bioinformatics Guides
- RNA-Seq Visualization: Volcano Plots, Heatmaps, and PCA
- Volcano Plot Proteomics: How to Create and Interpret Them Effectively
- RNA Sequencing Data Analysis: From Raw Reads to Differential Expression
- Proteomics Data Analysis in R: A Practical Workflow for Differential Expression and Visualization
- RNA-Seq vs Microarray: Choosing the Right Gene Expression Profiling Platform
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
[1] [Long non-coding RNA RP11-342L8.2, derived from RNA sequencing and validated via RT-qPCR, is upregulated and correlates with disease severity in psoriasis patients.](https://pubmed.ncbi.nlm.nih.gov/35028895). Irish journal of medical science, 2022. [2] [Upregulated ribosome pathway plays a key role in HDAC4, improving the survival rate and biofunction of chondrocytes.](https://pubmed.ncbi.nlm.nih.gov/37414410). Bone & joint research, 2023. [3] [Inflated false discovery rate due to volcano plots: problem and solutions.](https://pubmed.ncbi.nlm.nih.gov/33758907). Briefings in bioinformatics, 2021. [4] [Bioconductor](https://bioconductor.org/). Bioconductor Project. [5] [Using ExpressAnalyst for Comprehensive Gene Expression Analysis in Model and Non-Model Organisms.](https://pubmed.ncbi.nlm.nih.gov/37929753). Current protocols, 2023. [6] [TCC-GUI: a Shiny-based application for differential expression analysis of RNA-Seq count data.](https://pubmed.ncbi.nlm.nih.gov/30867032). BMC research notes, 2019. [7] [The Carpentries Lessons](https://carpentries.org/lessons). The Carpentries. [8] [Identification of HBEGF+ fibroblasts in the remission of rheumatoid arthritis by integrating single-cell RNA sequencing datasets and bulk RNA sequencing datasets.](https://pubmed.ncbi.nlm.nih.gov/36068607). Arthritis research & therapy, 2022. [9] [Landscape of N<,sup>,6<,/sup>,-methyladenosine (m<,sup>,6<,/sup>,A) methylation in porcine cells identifies candidate genes associated with porcine epidemic diarrhea virus infection.](https://doi.org/10.3389/fmicb.2026.1828985). 2026. [10] [Loss of KLF15 expression characterizes proximal tubule injury in cisplatin-induced acute kidney injury: A multi-omics study.](https://doi.org/10.1016/j.crtox.2026.100300). 2026. [11] [BRD4 PROTAC degrader enhances fulvestrant sensitivity in ER+ breast cancer via super-enhancer associated GREB1.](https://doi.org/10.3389/fonc.2026.1828900). 2026. [12] [SF3B4 stabilizes SREBF1 via 3'UTR binding to drive hepatocellular carcinoma progression.](https://doi.org/10.3389/fonc.2026.1833357). 2026. [13] [NCBI Data Resources](https://www.ncbi.nlm.nih.gov/). National Center for Biotechnology Information. [14] [Galaxy Training Network](https://training.galaxyproject.org/). Galaxy Project. [15] [nf-core Documentation](https://nf-co.re/docs). nf-core. [16] [EMBL-EBI Training](https://www.ebi.ac.uk/training). European Bioinformatics Institute. [17] [Transcriptomic profiling reveals substrate- and shear stress-dependent maturation of human small intestinal epithelial cells.](https://doi.org/10.3389/fphar.2026.1706106). 2026.This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.