Bonferroni vs. Benjamini-Hochberg FDR

By Dr. Zubair Khalid, DVM, MS, PhD ·

Bonferroni vs. Benjamini-Hochberg FDR

Key Takeaways

  • Bonferroni Correction: This method strictly controls the Family-Wise Error Rate (FWER), ensuring the probability of at least one false positive across all tests remains below the alpha level. It is best suited for confirmatory studies with a small number of pre-specified hypotheses, such as validating a few candidate genes identified in prior screens, where a single false positive could invalidate the entire conclusion.
  • Benjamini-Hochberg (BH) FDR Control: This procedure controls the False Discovery Rate (FDR), meaning it limits the expected proportion of false positives among all declared significant tests. It is more powerful than Bonferroni and is ideal for exploratory research, such as high-throughput screening in transcriptomics or proteomics, where identifying potential candidates for further validation is the primary goal.
  • Power vs. Error Control Trade-off: Bonferroni is highly conservative, leading to lower statistical power and a higher risk of false negatives, especially with many tests. BH offers increased power by accepting a controlled proportion of false positives, making it more suitable for detecting subtle effects in large-scale analyses like genome-wide association studies (GWAS).
  • Assumptions and Limitations: Both methods assume independence or weak dependence among tests; significant correlations can impact BH procedure's performance. Neither method can correct for fundamental issues like poor study design, biased measurements, or selective reporting, and adjusted p-values should always be interpreted alongside effect sizes.
  • Contextual Application: The choice hinges on the research objective: confirmatory studies with high cost of false positives necessitate Bonferroni, while exploratory studies prioritizing discovery and accepting a small false positive rate benefit from BH, particularly when dealing with thousands or millions of tests.

Quick Answer

  • Choose Bonferroni when your study is confirmatory and a single false positive would invalidate a conclusion, because it strictly controls the family-wise error rate.
  • Choose Benjamini-Hochberg false discovery rate control when your study is exploratory and you accept a small proportion of false positives to preserve power for detecting real effects.
  • Both methods assume independent or weakly dependent tests, and neither corrects for poor study design, biased measurements, or selective reporting.

At a Glance

FeatureBonferroni CorrectionBenjamini-Hochberg FDR
Error metric controlledFamily-wise error rate (probability of at least one false positive)False discovery rate (expected proportion of false positives among rejected hypotheses)
Mathematical approachMultiply each p-value by the number of tests, or compare each p-value to alpha divided by the number of testsSort p-values, compare each to its rank times alpha divided by the total number of tests, then find the largest significant threshold
Typical research contextConfirmatory studies, clinical endpoints, regulatory submissions, genome-wide significance thresholdsExploratory screening, differential expression analysis, candidate gene discovery, ecological community surveys
Statistical powerLower power, especially with many tests, because the threshold becomes very strictHigher power, because the threshold adapts to the distribution of observed p-values
False positive behaviorControls the chance of any false positive across all testsControls the proportion of false positives only among tests you declare significant
InterpretationA significant result is strong evidence against the null hypothesisA significant result is a candidate for further validation, not a confirmed finding
Common software implementationp.adjust(p, method = "bonferroni") in R, multipletests in Python statsmodelsp.adjust(p, method = "BH") in R, multipletests in Python statsmodels

The Multiple Testing Problem in Life Science Research

Modern biological experiments routinely generate thousands of simultaneous statistical tests. A single RNA sequencing experiment can compare expression levels for tens of thousands of genes. A metabolomics panel can measure hundreds of compounds. A genome-wide association study can test millions of genetic variants. Each individual test carries its own probability of producing a false positive, and when you perform many tests, the chance that at least one of them is significant by chance alone grows rapidly.

Consider a simple example. If you perform 100 independent tests and use the conventional alpha level of 0.05 for each test, you expect about 5 false positives by chance alone. The probability that at least one test is falsely significant is not 5 percent but approximately 99.4 percent. This inflation of error is the multiple testing problem, and it is the reason why raw p-values from large-scale biological experiments cannot be interpreted in isolation.

The multiple testing problem is not a mathematical curiosity. It has direct consequences for the reproducibility of biological research. When researchers report findings that are actually chance fluctuations, subsequent validation experiments fail, resources are wasted, and the scientific literature accumulates results that cannot be replicated. The Committee on Publication Ethics Core Practices emphasize that research should be conducted rigorously and reported honestly, and appropriate statistical correction is part of that rigor.

The two most common solutions to the multiple testing problem are the Bonferroni correction and the Benjamini-Hochberg procedure. Both adjust p-values to account for the number of tests performed, but they control different error metrics and answer different scientific questions. Understanding the distinction is essential for choosing the right method for your specific research context.

The Bonferroni Correction

The Bonferroni correction is the simplest and most conservative approach to multiple testing. It controls the family-wise error rate, which is the probability of making at least one false positive among all tests performed. The method sets a new significance threshold by dividing the desired alpha level by the number of tests.

If you perform 1,000 tests and want to keep the family-wise error rate at 0.05, the Bonferroni threshold is 0.05 divided by 1,000, which equals 0.00005. A test is declared significant only if its p-value is below this adjusted threshold. Equivalently, you can multiply each raw p-value by the number of tests and compare the adjusted value to the original alpha.

The Bonferroni correction is guaranteed to control the family-wise error rate regardless of how many tests you perform or how the tests are correlated. This guarantee is its main strength. It is also its main weakness, because the guarantee comes at the cost of statistical power. When the number of tests is large, the threshold becomes so strict that genuine but modest effects may not reach significance.

The correction is most appropriate when you have a small number of tests and when a single false positive would be scientifically damaging. For example, if you are testing a small set of candidate genes that were selected based on prior biological knowledge, the Bonferroni correction provides strong protection against claiming an effect that does not exist. It is also appropriate for confirmatory analyses where the study was designed to test a specific hypothesis and where the results will guide a consequential decision.

When Bonferroni Is Too Conservative

The Bonferroni correction assumes that all tests are independent. In biological data, tests are often correlated. Genes in the same pathway are co-expressed, metabolites in the same biochemical pathway are correlated, and SNPs in linkage disequilibrium are inherited together. When tests are correlated, the Bonferroni correction is overly strict because it treats each test as if it were completely independent.

The result is an increased rate of false negatives. A false negative occurs when a real effect exists but the statistical test fails to detect it. In a screening experiment where you are trying to identify candidate genes for further study, a false negative means you miss a gene that is genuinely involved in the biological process. This can be as damaging as a false positive, because it means you fail to pursue a promising lead.

The Bonferroni correction also becomes impractical when the number of tests is very large. In a genome-wide association study with millions of tests, the Bonferroni threshold is so small that only effects of very large magnitude can be detected. This is why the method is rarely used as the sole correction in high-throughput screening experiments.

The Benjamini-Hochberg Procedure

The Benjamini-Hochberg procedure, often called the BH procedure or FDR control, takes a different approach. Instead of controlling the probability of any false positive, it controls the false discovery rate, which is the expected proportion of false positives among all tests that are declared significant.

The procedure works by sorting all p-values from smallest to largest. Each p-value is then compared to a threshold that depends on its rank in the sorted list. The threshold for the kth smallest p-value is k times alpha divided by the total number of tests. The procedure finds the largest k such that the kth smallest p-value is below this threshold, and declares all tests with p-values smaller than that value as significant.

The key difference from Bonferroni is that the BH procedure allows a controlled proportion of false positives. If you set the false discovery rate at 0.05, you are accepting that about 5 percent of the tests you declare significant may be false positives. This is a different scientific statement than saying that the probability of any false positive is 5 percent.

The BH procedure is more powerful than Bonferroni because it does not apply the same strict threshold to every test. The smallest p-values are compared to a threshold that is close to the Bonferroni threshold, but larger p-values are compared to progressively more lenient thresholds. This means that the procedure can detect more true effects while still controlling the overall proportion of false discoveries.

When FDR Control Is Appropriate

The BH procedure is designed for exploratory research where the goal is to identify candidates for further investigation. In transcriptomics, proteomics, metabolomics, and other high-throughput screening experiments, the goal is often to generate hypotheses instead of to confirm them. A small proportion of false positives is acceptable because the findings will be validated in follow-up experiments.

The procedure is also appropriate when the number of tests is very large and the Bonferroni correction would be too strict to detect any effects. In a differential expression analysis with 20,000 genes, the BH procedure can identify a set of candidate genes for further study, while the Bonferroni correction might identify only the most extreme outliers.

The BH procedure assumes that the tests are independent or positively correlated. This assumption is often approximately satisfied in biological data, but it is worth checking. If the tests are negatively correlated, the procedure can be too liberal and the false discovery rate may exceed the nominal level.

Decision Framework for Choosing a Correction Method

The choice between Bonferroni and BH depends on the scientific goal of your analysis. The first question to ask is whether your study is confirmatory or exploratory.

A confirmatory study tests a small number of pre-specified hypotheses. The hypotheses were defined before the data were collected, and the results will be used to make a decision. In this context, a false positive is a serious error because it can lead to a wrong conclusion. The Bonferroni correction is the appropriate choice because it strictly controls the family-wise error rate.

An exploratory study screens many variables to generate hypotheses. The goal is to identify candidates for further study, and the results will be validated in subsequent experiments. In this context, a small proportion of false positives is acceptable because the cost of missing a real effect is higher than the cost of following up on a false lead. The BH procedure is the appropriate choice because it controls the false discovery rate while preserving power.

The table below summarizes the decision framework.

Research ContextNumber of TestsCost of False PositiveCost of False NegativeRecommended Method
Confirmatory clinical endpointSmall (1 to 20)High, a wrong conclusion affects patient careModerateBonferroni
Candidate gene validationSmall to moderate (10 to 100)High, a false positive wastes validation resourcesModerateBonferroni
Transcriptomics screeningLarge (10,000 to 30,000)Low, findings will be validatedHigh, missing a real gene loses a leadBenjamini-Hochberg
Proteomics discoveryLarge (1,000 to 10,000)Low, findings will be validatedHigh, missing a real protein loses a leadBenjamini-Hochberg
Genome-wide association studyVery large (millions)Low, findings will be replicatedHigh, missing a real variant loses a leadBenjamini-Hochberg with additional filters

Practical Workflow for Applying Multiple Testing Corrections

The practical workflow for applying multiple testing corrections involves several steps that should be documented in your analysis plan. The National Institutes of Health Grants and Funding pages describe the expectations for rigorous experimental design and statistical analysis in funded research, and the Data Management and Sharing Policy requires that your analysis methods be described clearly enough for others to understand and reproduce.

Step 1: Define the Number of Tests

The number of tests is the total number of statistical comparisons you perform. This includes all genes, all proteins, all metabolites, or all SNPs that you test. It is important to count the tests correctly. If you filter out genes with low expression before testing, the number of tests is the number of genes that remain after filtering. If you test multiple contrasts, such as treatment versus control and high dose versus low dose, each contrast is a separate test.

Step 2: Choose the Error Metric

Decide whether you need to control the family-wise error rate or the false discovery rate. This decision should be made before you look at the data. If you decide after seeing the results, you risk choosing the method that gives you the most significant results, which is a form of p-hacking.

Step 3: Apply the Correction

In R, the p.adjust function applies both methods. The code for Bonferroni is p.adjust(p_values, method = "bonferroni") and for Benjamini-Hochberg is p.adjust(p_values, method = "BH"). In Python, the multipletests function from the statsmodels library provides both methods.

Step 4: Report the Adjusted Values

Report the adjusted p-values in your results section. State which method you used and why. The EQUATOR Network provides reporting guidelines for many study types, and these guidelines often specify how to report statistical corrections.

Step 5: Document the Analysis

Record the number of tests, the correction method, the alpha level, and the software version in your analysis documentation. This documentation is required for reproducibility and is part of the Data Management and Sharing Policy expectations.

Common Failure Patterns

Several common mistakes occur when researchers apply multiple testing corrections.

Using the Wrong Error Metric

The most common mistake is using the Bonferroni correction when the research question is exploratory. This leads to a high false negative rate and can cause researchers to miss real effects. Conversely, using the BH procedure in a confirmatory study can lead to a higher rate of false positives than the study design can tolerate.

Applying the Correction to the Wrong Number of Tests

Some researchers apply the correction only to the tests that were performed in the final analysis, ignoring tests that were performed during data exploration. If you tested 100 genes in a pilot analysis and then tested 10 genes in the final analysis, the number of tests should reflect the total number of comparisons that were made, beyond the final set.

Ignoring Correlation Between Tests

The BH procedure assumes that tests are independent or positively correlated. When tests are negatively correlated, the procedure can be too liberal. The Bonferroni correction is valid regardless of correlation, but it is conservative when tests are positively correlated.

Using Adjusted P-Values Without Context

An adjusted p-value is not a measure of effect size. A gene can have a highly significant adjusted p-value and a very small effect that is biologically irrelevant. Always report the effect size and the confidence interval alongside the adjusted p-value.

Failing to Pre-Specify the Analysis

The choice of correction method should be made before the data are analyzed. If you choose the method after seeing the results, you are at risk of selecting the method that produces the most favorable results. This is a form of selective reporting and is a violation of research integrity standards described in the Committee on Publication Ethics Core Practices.

Limitations of Both Methods

Both the Bonferroni and the BH procedures have limitations that are important to understand.

The Bonferroni Correction Is Conservative

The Bonferroni correction is guaranteed to control the family-wise error rate, but it does so at the cost of power. When the number of tests is large, the threshold becomes so strict that many real effects are missed. This is a particular problem in high-throughput experiments where the number of tests is in the thousands or millions.

The BH Procedure Does Not Control the Family-Wise Error Rate

The BH procedure controls the false discovery rate, but it does not control the probability of any false positive. If you have 1,000 tests and you declare 100 significant at a false discovery rate of 0.05, you expect about 5 of those 100 to be false positives. This is acceptable in exploratory research, but it is not acceptable in a confirmatory study where a single false positive can invalidate the conclusion.

Both Methods Assume the Tests Are Valid

Both methods assume that the statistical tests are valid. If the data violate the assumptions of the underlying test, such as normality or equal variance, the p-values are not reliable and the correction is not meaningful. The correction cannot fix a fundamentally flawed analysis.

Both Methods Assume the Data Are Complete

The correction methods assume that the data are complete and that the tests are performed on all the data. If you filter out tests based on the results, the correction is not valid. For example, if you perform a differential expression analysis and then remove genes with very low expression, the number of tests should be the number of genes that were tested, not the number that remained after filtering.

Welfare and Safety Context

The choice of multiple testing correction has direct consequences for the safety and welfare of research subjects and for the integrity of the research enterprise.

In a clinical trial, a false positive can lead to the approval of an ineffective treatment or the rejection of an effective one. The Bonferroni correction is appropriate in this context because the cost of a false positive is high. In a preclinical study, a false positive can lead to the pursuit of a drug target that is not real, wasting resources and potentially delaying the development of an effective treatment.

In animal research, the choice of correction method affects the number of animals needed. A more powerful method, such as the BH procedure, can detect effects with a smaller sample size, reducing the number of animals required. However, the increased power comes at the cost of a higher false positive rate, which can lead to the pursuit of false leads.

The National Institutes of Health Grants and Funding pages describe the requirements for rigorous experimental design and statistical analysis in funded research. The Data Management and Sharing Policy requires that the data and the analysis methods be shared in a way that allows others to reproduce the results.

Professional Escalation Criteria

There are situations where you should seek professional statistical advice before applying a multiple testing correction.

When the Number of Tests Is Very Large

When the number of tests is in the millions, as in a genome-wide association study, the standard BH procedure may not be appropriate. More advanced methods, such as the Benjamini-Hochberg procedure with a more stringent threshold or the Storey q-value method, may be needed. A professional statistician can help you choose the appropriate method.

When the Tests Are Strongly Correlated

When the tests are strongly correlated, such as when you are testing many genes in the same pathway, the BH procedure may not control the false discovery rate at the nominal level. A statistician can help you assess the correlation structure and choose an appropriate method.

When the Results Will Guide a Regulatory Decision

When the results of your analysis will be used to support a regulatory submission, such as a new drug application, you should consult a statistician who is familiar with the regulatory requirements. The National Institutes of Health Grants and Funding pages describe the statistical rigor expected in funded research, and the EQUATOR Network provides reporting guidelines that may be relevant.

When the Data Are Missing or Incomplete

When the data are missing or incomplete, the standard correction methods may not be valid. A statistician can help you use methods that account for missing data.

Records and Measurements

The choice of multiple testing correction should be documented in your analysis plan and in your final report. The documentation should include the following:

  • The number of tests performed
  • The correction method used
  • The alpha level used
  • The software and version used
  • The date of the analysis
  • The name of the person who performed the analysis

This documentation is part of the Data Management and Sharing Policy and is required for reproducibility.

Common Failure Patterns

The following failure patterns are common in the application of multiple testing corrections.

Failure to Apply Any Correction

The most common failure is to apply no correction at all. This is a serious error in any study that performs more than a handful of tests. The result is a high rate of false positives and a study that cannot be replicated.

Applying the Correction to the Wrong Number of Tests

Some researchers apply the correction to the number of tests that were performed in the final analysis, but not to the number of tests that were performed during data exploration. This is a form of selective reporting and can lead to a false positive rate that is higher than the nominal level.

Applying the Correction to the Wrong Type of Test

Some researchers apply the correction to the p-values from a test that is not appropriate for the data. For example, applying a correction to the p-values from a test that assumes normality when the data are not normal is not meaningful.

Applying the Correction to the Wrong Alpha Level

Some researchers apply the correction to an alpha level that is not appropriate for the research context. For example, using an alpha level of 0.05 in a study that requires a more stringent level, such as a clinical trial, is not appropriate.

Applying the Correction to the Wrong Data

Some researchers apply the correction to the data that are not the data that were used to generate the p-values. For example, applying the correction to the p-values from a subset of the data that was selected after seeing the results is a form of bias.

Limitations of the Correction Methods

The correction methods are not a substitute for good experimental design. A well-designed experiment with a small number of tests and a clear hypothesis is more informative than a poorly designed experiment with a large number of tests and a correction method.

The correction methods also do not address the problem of publication bias. If researchers only publish significant results, the literature will be biased toward false positives, regardless of the correction method used. The Committee on Publication Ethics Core Practices address the need for transparent reporting of all results, including null results.

The correction methods do not address the problem of p-hacking. If a researcher tries many different analyses and reports only the one that gives the most significant results, the correction method cannot fix the bias. The correction method is applied to the p-values from a single analysis, not to the p-values from all the analyses that were tried.

The Role of Reporting Guidelines

The EQUATOR Network provides reporting guidelines for many types of research. These guidelines often specify the statistical methods that should be reported, including the multiple testing correction method. Following these guidelines improves the transparency and reproducibility of your research.

The Committee on Publication Ethics Core Practices require that research be reported honestly and transparently. This includes reporting the multiple testing correction method and the number of tests that were performed.

The Role of Data Management

The Data Management and Sharing Policy requires that the data and the analysis methods be shared in a way that allows others to reproduce the results. This includes the multiple testing correction method and the number of tests that were performed.

The ORCID for Researchers pages describe the importance of maintaining a record of your research outputs. This record includes the analysis methods that you used.

A Field Log System for Tracking Correction Decisions and Their Consequences

The choice between Bonferroni and Benjamini-Hochberg is not a one-time decision that ends when you run the p.adjust function. In practice, the consequences of that choice unfold over the entire research project, from the first exploratory screen to the final confirmatory experiment. A field log system that records your correction decisions, the rationale behind them, and the downstream outcomes provides a concrete way to evaluate whether your initial choice was appropriate. This system is especially valuable in multi-stage research projects where you apply different corrections at different phases.

Why a Written Decision Log Matters

The Committee on Publication Ethics Core Practices emphasize that research should be conducted rigorously and reported honestly. A decision log is a practical tool for meeting that expectation. It creates a contemporaneous record of what you decided, when you decided it, and why. This record serves three distinct purposes.

First, it prevents post-hoc rationalization. When you record your correction choice before you see the results, you cannot later claim that you always planned to use the BH procedure when the Bonferroni correction would have made your results non-significant. The log makes your pre-specified plan visible and auditable.

Second, it supports reproducibility. The Data Management and Sharing Policy requires that your analysis methods be described clearly enough for others to understand and reproduce. A decision log that records the number of tests, the correction method, the alpha level, and the date of the decision is a concrete way to meet that requirement.

Third, it creates a learning record. When you look back at a completed study, the log tells you whether your correction choice produced the expected balance of false positives and false negatives. This information helps you calibrate your decisions for future studies.

The Decision Log Template

A practical decision log has a simple structure. Each entry records one correction decision. The entry should be completed before you run the analysis, not after. The following fields are essential.

FieldWhat to RecordExample Entry
Study identifierThe name or code for the studyRNAseq_2024_heatstress
Analysis stageExploratory screen, candidate validation, or confirmatory testExploratory screen
Date of decisionThe date you decided on the method2026-03-14
Number of testsThe total number of statistical tests planned18,432 genes
Correction methodBonferroni or Benjamini-HochbergBenjamini-Hochberg
Alpha levelThe nominal error rate you are controlling0.05
RationaleThe scientific reason for the choiceExploratory screen, findings will be validated in qPCR
Expected consequenceWhat you expect in terms of false positives and false negativesExpect about 5 percent false positives among significant genes, accept some false negatives

The rationale field is the most important. It forces you to state the scientific context that drives your choice. A rationale that says "because the BH procedure gives more significant results" is not acceptable. A rationale that says "because this is an exploratory screen and the findings will be validated in a separate confirmatory experiment" is a legitimate scientific justification.

Recording the Consequences of the Decision

The decision log does not end when you run the analysis. You should also record the consequences of your decision. This is the part that most researchers skip, and it is the part that provides the most useful information for future studies.

After you run the analysis, record the following:

  • The number of tests that were significant after correction
  • The number of tests that were significant before correction but not after correction
  • The range of adjusted p-values for the significant tests
  • The effect sizes for the significant tests
  • Any validation results that you have obtained

This consequence record is what allows you to evaluate your decision. If you used the Bonferroni correction and found no significant results, the log tells you whether the correction was too strict for the effect sizes in your data. If you used the BH procedure and found many significant results, the log tells you whether the proportion of false positives was acceptable in the context of your validation results.

A Worked Example from a Transcriptomics Screen

Consider a transcriptomics experiment that compares gene expression between a treatment group and a control group. The experiment measures 18,432 genes. The researcher records the following decision log entry before running the analysis.

Study identifier: RNA_2026_heatstress Analysis stage: Exploratory screen Date of decision: 2026-03-14 Number of tests: 18,432 Correction method: Benjamini-Hochberg Alpha level: 0.05 Rationale: Exploratory screening, findings will be validated in follow-up qPCR Cost of false positive: Low, because findings will be validated Cost of false negative: High, because missing a real gene loses a lead

The researcher runs the analysis and finds that 1,204 genes are significant at a false discovery rate of 0.05. The researcher records the consequences in the log.

Number of significant tests after correction: 1,204 Number of significant tests before correction: 2,847 Proportion of significant tests after correction: 6.5 percent Effect sizes for significant tests: Log2 fold change range from 0.4 to 3.2 Follow-up validation: 40 genes selected for qPCR validation, 32 confirmed

The consequence record shows that the BH procedure identified a manageable set of candidate genes and that the validation rate was 80 percent. This information is useful for the next study. If the researcher runs a similar experiment in the future, the log provides a basis for expecting a similar proportion of significant results and a similar validation rate.

Now consider the same experiment with a Bonferroni correction. The researcher records the decision in the log.

Study identifier: RNA_2026_heatstress Analysis stage: Exploratory screen Date of decision: 2026-03-14 Number of tests: 18,432 Correction method: Bonferroni Alpha level: 0.05 Rationale: To avoid any false positives in the screen Cost of false positive: Low, findings will be validated Cost of false negative: High, missing a real gene loses a lead

The analysis produces 87 significant genes. The researcher records the consequences.

Number of significant tests after correction: 87 Number of significant tests before correction: 2,847 Proportion of significant tests after correction: 0.5 percent Effect sizes for significant tests: Log2 ratios range from 1.8 to 3.2 Follow-up validation: 40 of the significant genes tested by qPCR, 35 percent confirmed

The Bonferroni correction produced a smaller set of significant genes, and the effect sizes were larger. The validation rate was higher, but the researcher missed many genes with smaller effect sizes that might have been biologically relevant. The decision log shows that the Bonferroni correction was not appropriate for this exploratory screen because the cost of missing a real gene was high and the researcher accepted a high false negative rate.

Using the Log to Compare Methods Across Studies

The decision log becomes more powerful when you use it across multiple studies. If you maintain a log for all of your analyses, you can compare the consequences of different correction decisions across similar experiments. This comparison helps you calibrate your choices for future work.

For example, suppose you have three transcriptomics studies in your log. In the first study, you used the Bonferroni correction and found 87 significant genes. In the second study, you used the BH procedure and found 1,204 significant genes. In the third study, you used the BH procedure and found 2,100 significant genes. The log shows that the BH procedure consistently produces a larger set of significant genes, and the validation rates in the follow-up experiments tell you whether the larger set is mostly true positives or mostly false positives.

This comparison is not a statistical test. It is a practical record that helps you understand the behavior of the correction methods in your specific research context. The National Library of Medicine Research Methods Resources provide a gateway to authoritative biomedical books and research-method references that can help you interpret the patterns you see in your log.

Troubleshooting When the Log Reveals a Problem

The decision log can reveal problems that are not visible when you look at a single analysis. The following patterns indicate that your correction choice may need to be revisited.

Pattern 1: The BH Procedure Produces a Very Large Proportion of Significant Tests

If the BH procedure declares more than 20 percent of your tests significant, the procedure may be too liberal for your data. This can happen when the tests are strongly correlated, when the data contain technical artifacts, or when the underlying effect is very large. The log should prompt you to examine the distribution of raw p-values. If the p-values are not uniformly distributed under the null, the BH procedure may not control the false discovery rate at the nominal level.

Pattern 2: The Bonferroni Correction Produces No Significant Tests

If the Bonferroni correction produces no significant tests in a study where you expected to find effects, the correction may be too strict for your data. The log tells you the number of tests and the alpha level, so you can calculate the Bonferroni threshold and compare it to the distribution of raw p-values. If the smallest p-values are close to the threshold but not below it, the study may be underpowered or the effect sizes may be smaller than expected.

Pattern 3: The Validation Rate Is Very Low

If you used the BH procedure and the follow-up validation rate is below 10 percent, the false discovery rate may be higher than the nominal level. The log records the alpha level and the number of significant tests. You can use this information to decide whether to use a more stringent threshold in the next study or to apply a different correction method.

Pattern 4: The Results Change Substantially Between Stages

If you use the BH procedure in an exploratory screen and the Bonferroni correction in a confirmatory study, the log shows whether the results are consistent between the stages. If many of the genes that were significant in the exploratory screen are not significant in the confirmatory study, the log tells you that the exploratory screen produced a high proportion of false positives. This information is useful for deciding whether to use a stricter threshold in the next exploratory screen.

The Log as a Pre-Registration Tool

The decision log can serve as a simple form of pre-registration for your analysis. Pre-registration is the practice of recording your analysis plan before you collect or analyze the data. The Committee on Publication Ethics Core Practices describe the importance of transparent research practices, and pre-registration is one way to demonstrate that your analysis was planned in advance.

The log does not need to be a formal pre-registration document. A simple spreadsheet with the fields described above is sufficient. The key is that the decision is recorded before the analysis is run, and the log is stored in a place where it cannot be altered after the fact. This can be a dated file on a shared drive, a laboratory notebook, or a version-controlled repository.

The Log and the Data Management Plan

The NIH Data Management and Sharing Policy requires that your data and analysis methods be shared in a way that allows others to reproduce the results. The decision log is a natural part of this requirement. When you share your data, you can include the decision log as a companion document that explains the correction decisions.

The log also supports the ORCID for Researchers record. When you publish a paper, you can include the decision log as supplementary material. The log provides a transparent record of the analysis decisions, and it can be linked to your ORCID record to demonstrate your research practices.

A Simple Log Template

The following template can be copied into a spreadsheet or a text file. Each row is one analysis.

Study IDStageDateNumber of TestsMethodAlphaRationaleSignificant After CorrectionSignificant Before CorrectionValidation Rate
RNA_2026_heatExploratory2026-03-1418,432BH0.05Screening, will validate1,2042,84732 percent
RNA_2026_heatConfirmatory2026-05-0240Bonferroni0.05Confirmatory qPCR1315Not applicable
Protein_2026_kinaseExploratory2026-04-104,500BH0.05Screening, will validate31289018 percent

The template is simple, but it captures the information that is needed to evaluate the correction decision. The validation rate column is the most informative because it tells you whether the correction method produced results that were confirmed in follow-up experiments.

When to Escalate to a Professional Statistician

The decision log can also help you identify when you need professional statistical advice. The National Institutes of Health Grants and Funding pages describe the statistical rigor expected in funded research, and the log can help you identify situations where the standard methods are not sufficient.

If the log shows that the BH procedure consistently produces a high proportion of significant tests with a low validation rate, you should consult a statistician. The statistician can help you assess whether the data violate the assumptions of the BH procedure and whether a more advanced method is needed.

If the log shows that the Bonferroni correction consistently produces no significant tests in studies where you expect effects, you should consult a statistician. The statistician can help you assess the power of your study and whether the correction is appropriate for your research context.

If the log shows that the results change substantially between the exploratory and confirmatory stages, you should consult a statistician. The statistician can help you understand the sources of the discrepancy and whether the correction methods are being applied correctly.

The Log and the Reporting Guidelines

The EQUATOR Network provides reporting guidelines for many study types. These guidelines often specify the statistical methods that should be reported, including the multiple testing correction method. The decision log is a practical tool that helps you report the required information.

When you write your methods section, you can use the log to state the number of tests, the correction method, the alpha level, and the rationale for the choice. When you write your results section, you can use the log to report the number of significant tests and the adjusted p-values. The log makes the reporting process straightforward because the information is already recorded.

The Log as a Learning Tool

The decision log is a learning tool that improves your research practice over time. Each entry in the log is a record of a decision and its consequences. When you review the log at the end of a study or at the end of a year, you can see patterns in your decisions and their outcomes.

You may find that you consistently use the BH procedure in exploratory screens and that the validation rates are consistently above 20 percent. This pattern tells you that the BH procedure is working well in your research context. You may find that you consistently use the Bonferroni correction in confirmatory studies and that the results are consistently reproducible. This pattern tells you that the Bonferroni correction is appropriate for your confirmatory work.

You may also find that you are using the Bonferroni correction in exploratory screens and that the validation rates are high but the number of significant results is very small. This pattern tells you that the Bonferroni correction is too strict for your exploratory work and that you are missing real effects. The log provides the evidence you need to change your practice.

The Log and the Research Record

The decision log is a simple tool that supports the integrity of your research. It records your decisions, the rationale behind them, and the consequences of those decisions. It is a practical way to implement the expectations described in the Committee on Publication Ethics Core Practices and the NIH Data Management and Sharing Policy.

The log is not a substitute for a professional statistician. It is a tool that helps you make better decisions and to document those decisions in a way that is transparent and auditable. When you use the log consistently, you build a record of your research practice that is valuable for your own learning and for the reproducibility of your work.

Frequently Asked Questions

What is the difference between the Bonferroni and the Benjamini-Hochberg methods?

The Bonferroni method controls the family-wise error rate, which is the probability of at least one false positive across all tests. The Benjamini-Hochberg method controls the false discovery rate, which is the expected proportion of false positives among the tests that are declared significant. The Bonferroni method is more conservative and is appropriate for confirmatory studies. The Benjamini-Hochberg method is more powerful and is appropriate for exploratory studies.

When should I use the Bonferroni correction?

Use the Bonferroni correction when you have a small number of tests and when a single false positive would invalidate your conclusion. This is common in confirmatory studies, such as clinical trials or candidate gene validation studies.

When should I use the Benjamini-Hochberg method?

Use the Benjamini-Hochberg method when you have a large number of tests and when you are screening for candidates for further study. This is common in exploratory studies, such as transcriptomics or proteomics screening.

Can I use both methods in the same study?

You can use both methods in the same study, but you should use them for different purposes. For example, you might use the Bonferroni method for the primary confirmatory analysis and the Benjamini-Hochberg method for the secondary exploratory analysis. The methods should be pre-specified in the analysis plan.

What is the false discovery rate?

The false discovery rate is the expected proportion of false positives among the tests that are declared significant. If you set the false discovery rate at 0.05, you expect that about 5 percent of the significant results are false positives.

What is the family-wise error rate?

The family-wise error rate is the probability of at least one false positive across all tests. If you set the family-wise error rate at 0.05, the probability that any of the tests is a false positive is 5%.

How do I report the multiple testing correction in my paper?

Report the correction method, the number of tests, and the alpha level in the methods section. Report the adjusted p-values in the results section. The EQUATOR Network provides reporting guidelines that specify the details.

What should I do if the results are not significant after the correction?

If the results are not significant after the correction, you should report the results as not significant. You should not report the uncorrected p-values as significant. You can also consider whether the study was underpowered and whether a larger sample size is needed.

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.