Reporting Effect Sizes in Biological Papers: A Template with Examples
By Dr. Zubair Khalid, DVM, MS, PhD ·

Key Takeaways
- Report effect sizes with confidence intervals in all results sections, alongside p-values, to quantify the magnitude of biological findings and facilitate meta-analyses. Select the appropriate effect size metric (e.g., Cohen's d for two-group continuous outcomes, odds ratios for categorical outcomes, Pearson's r for continuous associations) before data collection, documenting this choice and the statistical software used in the methods section.
- Effect sizes are crucial for distinguishing statistical significance from biological significance; a statistically significant p-value does not inherently imply a biologically meaningful effect, and effect sizes provide this critical context. For instance, a Cohen's d of 0.2 might be biologically trivial in some contexts but significant in others, necessitating interpretation within the specific biological system.
- Common effect size families include the 'd' family (e.g., Cohen's d, Hedges' g) for standardized mean differences, the 'r' family (e.g., Pearson's r) for correlations, odds ratios for binary outcomes, and variance-explained measures (e.g., eta-squared) for ANOVA designs. The choice depends on the outcome variable's scale (continuous, ordinal, categorical) and the predictor variable's structure (categorical or continuous).
- Interpretation of effect sizes is context-dependent and field-specific; universal thresholds for "small," "medium," or "large" effects are not applicable across all biological disciplines. For example, a correlation of 0.3 might be considered moderate in genomics but large in studies of complex behavioral traits.
- A structured decision framework involving three questions: outcome scale, predictor structure, and primary research question, guides the selection of the appropriate effect size metric. This ensures the chosen metric aligns with the study design and research question, preventing common errors like using default software outputs or mismatched metrics for association versus difference questions.
Quick Answer
- Report effect sizes with confidence intervals in every results section, beyond p values, to support biological interpretation and reproducibility.
- Select the effect size metric before data collection and state it in the methods section with the statistical software used.
- A limitation is that effect size interpretation depends on the biological context and field-specific benchmarks, so universal thresholds do not apply.
Why Effect Size Reporting Matters in Biological Research
Biological research depends on quantifying the magnitude of observed phenomena. A statistically significant p value indicates that an observed pattern is unlikely under a null hypothesis, but it does not communicate the size or biological importance of that pattern. Two studies can both report p values below 0.05 while one shows a trivial difference and the other shows a substantial shift in a biological process. Effect sizes address this gap by quantifying the strength of an association or the magnitude of a difference between groups.
The reproducibility crisis in life sciences has drawn attention to incomplete statistical reporting. When manuscripts omit effect sizes, readers cannot assess whether a finding has practical biological relevance, cannot compare results across studies, and cannot include the result in future meta-analyses. The EQUATOR Network maintains a comprehensive collection of reporting guidelines that help authors select the appropriate standards for their study design. These guidelines consistently emphasize the need for effect size reporting alongside confidence intervals.
For laboratory professionals and biology students, understanding effect size reporting is a core skill. It affects how you interpret your own data, how you evaluate the literature, and how you design studies with adequate statistical power. The National Library of Medicine Research Methods Resources provides access to authoritative biomedical texts that cover the statistical foundations of effect size estimation and interpretation.
Core Principles of Effect Size Reporting
What Effect Sizes Communicate
An effect size is a standardized or unstandardized measure that quantifies the magnitude of a phenomenon. In biological research, effect sizes serve three primary functions. First, they provide a scale for interpreting the practical importance of a finding. Second, they enable comparison of results across studies that may use different measurement scales. Third, they are essential inputs for power analysis and sample size calculation in future studies.
The distinction between statistical significance and biological significance is central to effect size reporting. A large sample can produce a statistically significant p value for a biologically negligible effect. Conversely, a small sample may fail to reach statistical significance for a biologically meaningful effect. Effect sizes help researchers and readers distinguish these scenarios.
Common Effect Size Families in Biology
Biological research uses several families of effect size measures depending on the study design and data type.
The d family includes standardized mean differences such as Cohen's d, Hedges' g, and Glass's delta. These measures express the difference between two group means in units of standard deviation. They are appropriate for comparing two groups on a continuous outcome, such as comparing gene expression levels between treatment and control conditions.
The r family includes correlation coefficients and measures of association strength. Pearson's r is used for linear relationships between two continuous variables. Spearman's rho is used for monotonic relationships that may not be linear. These measures are common in genomics and proteomics where researchers examine relationships between molecular measurements.
The odds ratio family is used for binary outcomes. Odds ratios express the odds of an event in one group relative to another. They are common in genetic association studies and clinical research where outcomes are categorical.
The variance-explained family includes eta squared, partial eta squared, and omega squared. These are used in analysis of variance designs to express the proportion of variance in the outcome that is attributable to the grouping variable.
The Relationship Between Effect Size and Power
Statistical power is the probability of detecting an effect when one truly exists. Power depends on the sample size, the significance threshold, and the effect size. Researchers who plan studies without considering effect sizes risk designing studies that are underpowered to detect biologically meaningful effects.
The NIH Grants and Funding resources describe the expectations for rigorous experimental design in grant applications. Reviewers evaluate whether the proposed sample size is justified by the expected effect size and the statistical power of the planned analyses. A grant application that specifies the expected effect size and the power calculation demonstrates a stronger understanding of the study design than one that omits these details.
The Effect Size Reporting Template
The template below provides a structured approach to reporting effect sizes in biological manuscripts. It is designed to be adapted to different study types and statistical methods.
Template for Methods Section
The methods section should describe the effect size reporting plan before the results are presented. Include the following elements:
- State the primary outcome variables and the effect size metric that will be used for each.
- Justify the choice of effect size metric based on the study design and the nature of the data.
- Describe the statistical software and version used for all analyses.
- State the confidence interval level that will be reported with each effect size.
- Describe how missing data will be handled in the effect size calculations.
- State the threshold for interpreting the effect size if a field-specific convention exists.
Template for Results Section
The results section should report effect sizes for every primary analysis. Include the following elements:
- Report the effect size estimate with its confidence interval for each primary comparison.
- Report the p value alongside the effect size, but do not substitute one for the other.
- Describe the direction of the effect in biological terms.
- Report the effect size for secondary analyses when they support the primary findings.
- Include the effect size in figures and tables where appropriate.
Template for Discussion Section
The discussion section should interpret the effect sizes in the context of the biological question. Include the following elements:
- Interpret the magnitude of the effect size in biological terms.
- Compare the effect size with those reported in similar studies.
- Discuss the clinical or biological importance of the effect size.
- Acknowledge the uncertainty in the effect size estimate as reflected in the confidence interval.
Effect Size Metrics by Study Design
Two-Group Comparisons
For studies comparing two independent groups on a continuous outcome, the standardized mean difference is the most common effect size metric. Cohen's d is calculated as the difference between the two group means divided by the pooled standard deviation. Hedges' g includes a correction for small sample sizes and is preferred when sample sizes are below 20 per group.
The interpretation of the standardized mean difference depends on the biological context. A standardized difference of 0.2 is often considered small, 0.5 is considered medium, and 0.8 is considered large. These thresholds are conventions and should be interpreted with caution. In some biological systems, a standardized difference of 0.2 may be biologically important, while in others a difference of 1.0 may be trivial.
Correlation and Association Studies
For studies examining the relationship between two continuous variables, the correlation coefficient is the appropriate effect size. Pearson's r is appropriate when the relationship is linear and the variables are approximately normally distributed. Spearman's rho is appropriate when the relationship is monotonic but not necessarily linear.
The correlation coefficient is bounded between -1 and 1. The absolute value indicates the strength of the association, and the sign indicates the direction. A correlation of 0.3 is often considered a moderate effect, while a correlation of 0.5 is considered large in many biological contexts.
Categorical Outcomes
For studies with categorical outcomes, the odds ratio is a common effect size metric. The odds ratio expresses the odds of an outcome in one group relative to the odds in another group. An odds ratio of 1 indicates no association, an odds ratio greater than 1 indicates increased odds, and an odds ratio less than 1 indicates decreased odds.
Odds ratios are often reported with 95% confidence intervals. The interpretation of the odds ratio magnitude depends on the baseline risk of the outcome. An odds ratio of 2.0 may be substantial for a common outcome but less meaningful for a rare outcome.
Analysis of Variance Designs
For studies with more than two groups, the effect size is often expressed as eta squared or partial eta squared. These measures indicate the proportion of variance in the outcome that is explained by the grouping variable. Eta squared values range from 0 to 1, with higher values indicating a stronger effect.
Omega squared is a less biased alternative to eta squared, particularly when sample sizes are small. Researchers should report the effect size measure that is appropriate for their design and should state which measure they used.
At a Glance
| Study Design | Recommended Effect Size Metric | Interpretation Context | Reporting Requirement |
|---|---|---|---|
| Two independent groups, continuous outcome | Cohen's d or Hedges' g | Standardized mean difference in standard deviation units | Report with 95% confidence interval |
| Correlation between two continuous variables | Pearson's r or Spearman's rho | Strength and direction of association | Report with 95% confidence interval |
| Categorical outcome, two groups | Odds ratio | Relative odds of the outcome | Report with 95% confidence interval |
| Analysis of variance with multiple groups | Partial eta squared or omega squared | Proportion of variance explained | Report with confidence interval when available |
Practical Implementation Steps
Step 1: Select the Effect Size Metric Before Data Collection
The effect size metric should be selected during the study design phase, not after the data are collected. The choice of metric depends on the study design, the type of data, and the primary research question. Document the choice in the study protocol or the methods section of the manuscript.
Step 2: Calculate the Effect Size and Confidence Interval
Use the statistical software that is appropriate for the analysis. Most statistical packages provide effect size calculations for common analyses. For the analyses that do not provide effect sizes directly, use the formulas or dedicated effect size calculators.
Step 3: Report the Effect Size in the Results Section
Report the effect size estimate with its confidence interval for every primary analysis. Include the effect size in the text, in tables, and in figures where appropriate. Do not report the p-value without the effect size.
Step 4: Interpret the Effect Size in the Discussion Section
Interpret the effect size in the context of the biological question. Compare the effect size with those reported in similar studies. State whether the effect size is biologically meaningful.
Step 5: Deposit the Data and Analysis Code
The NIH Data Management and Sharing Policy describes the expectations for data sharing in NIH-funded research. Sharing the data and the analysis code allows other researchers to verify the effect size calculations and to include the results in future meta-analyses.
Records and Measurements
What to Record in the Laboratory Notebook
The laboratory notebook should include the following information related to effect size reporting:
- The effect size metric selected for each primary analysis.
- The statistical software and version used for the calculations.
- The sample size for each group in the analysis.
- The effect size estimate and its confidence interval.
- The p-value for each primary analysis.
- Any assumptions that were checked before the effect size calculation.
What to Record in the Manuscript
The manuscript should include the following information related to effect size reporting:
- The effect size metric in the methods section.
- The effect size estimate and confidence interval in the results section.
- The interpretation of the effect size in the discussion section.
- The data availability statement that describes where the data and analysis code can be accessed.
Common Failure Patterns in Effect Size Reporting
Reporting Only P-Values
The most common failure is reporting only p-values without effect sizes. This pattern prevents readers from assessing the biological importance of the findings. It also prevents the results from being included in future meta-analyses.
Reporting Effect Sizes Without Confidence Intervals
An effect size without a confidence interval does not convey the uncertainty in the estimate. A large effect size with a wide confidence interval may not be a reliable estimate. The confidence interval should always be reported with the effect size.
Using the Wrong Effect Size Metric
The effect size metric must match the study design and the data type. Using a correlation coefficient for a categorical outcome or an odds ratio for a continuous outcome produces misleading results.
Misinterpreting the Magnitude of the Effect Size
The conventional thresholds for small, medium, and large effects are not universal. The interpretation of the effect size must be based on the biological context and the field of study.
Failing to Report the Effect Size for Secondary Analyses
Secondary analyses should also report effect sizes. The effect size for a secondary analysis may be important for the interpretation of the primary findings.
Quality Controls and Reproducibility
Verification of Effect Size Calculations
The effect size calculations should be verified by a second researcher or by using a different software package. The verification should be documented in the laboratory notebook.
Reproducibility of the Analysis
The analysis should be reproducible by other researchers. This requires the data and the analysis code to be available. The NIH Data Management and Sharing Policy describes the expectations for data sharing in NIH-funded research.
Reporting Guidelines
The EQUATOR Network provides a comprehensive collection of reporting guidelines for different study designs. Authors should select the appropriate guideline for their study design and follow the reporting requirements, which include effect size reporting.
Limitations and Interpretation
The Effect Size Is an Estimate
The effect size is an estimate based on the sample data. The confidence interval provides a range of plausible values for the true effect size. The width of the confidence interval depends on the sample size and the variability of the data.
The Effect Size Does Not Establish Causation
An effect size from an observational study does not establish causation. The effect size indicates the magnitude of the association, but the causal interpretation requires additional evidence from the study design.
The Effect Size Depends on the Study Population
The effect size may vary across different study populations. The effect size reported in one study may not be generalizable to other populations.
The Effect Size Depends on the Measurement Method
The effect size depends on the measurement method used in the study. Different measurement methods may produce different effect sizes for the same biological phenomenon.
Safety and Regulatory Context
Reporting Standards in Peer-Reviewed Journals
The Committee on Publication Ethics Core Practices describe the expectations for ethical conduct in research publication. These practices include the responsible reporting of research findings, which includes the reporting of effect sizes.
Data Sharing Requirements
The NIH Data Management and Sharing Policy describes the expectations for data sharing in NIH-funded research. The policy requires that data be shared in a way that is consistent with the principles of reproducibility.
Researcher Identity and Record Keeping
The ORCID for Researchers resource describes the importance of maintaining a persistent researcher identifier. The ORCID record can be used to link the researcher to their publications and data sets, which supports the transparency of the research record.
Professional Escalation Criteria
When to Consult a Biostatistician
Consult a biostatistician when the study design is complex, when the data do not meet the assumptions of the planned statistical tests, or when the effect size metric is not clear. A biostatistician can provide guidance on the appropriate effect size metric and the interpretation of the results.
When to Consult a Data Manager
Consult a data manager when the data are complex, when the data are not well organized, or when the data sharing requirements are not clear. A data manager can help with the data organization and the data sharing plan.
When to Consult a Research Integrity Officer
Consult a research integrity officer when there is a question about the responsible conduct of research, including the reporting of data and the handling of conflicts of interest. The Committee on Publication Ethics Core Practices provide guidance on the responsible conduct of research.
A Decision Framework for Selecting and Defending Effect Size Metrics
Selecting an effect size metric is not a purely statistical decision. It is a biological decision that shapes how readers interpret your findings and whether your results can be combined with other studies. Many manuscripts fail at this step because the authors choose a metric based on what their software outputs by default instead of what the study design and research question require. This section provides a structured decision framework that you can apply before data collection, a record system for documenting your choices, and a troubleshooting method for common problems that arise during the selection and reporting process.
The Three Question Filter for Metric Selection
Before you open any statistical software, answer three questions about each primary outcome. Write the answers in your study protocol or laboratory notebook. The answers determine which effect size family is appropriate.
Question 1: What is the scale of the outcome variable?
The outcome variable is either continuous, ordinal, or categorical. Continuous variables include measurements such as enzyme activity, body mass, gene expression level, or growth rate. Ordinal variables include ranked categories such as disease severity scores or histological grades. Categorical variables include binary outcomes such as survival or death, presence or absence of a mutation, and nominal categories such as species identity.
The scale of the outcome determines the broad family of effect size metrics. Continuous outcomes support standardized mean differences and correlation coefficients. Ordinal outcomes support rank-based correlation coefficients such as Spearman's rho. Categorical outcomes support odds ratios, risk ratios, and risk differences.
Question 2: What is the structure of the predictor variable?
The predictor variable is either categorical or continuous. A categorical predictor has two or more groups, such as treatment versus control, or four different doses. A continuous predictor is a measured quantity such as concentration, time, or temperature.
The combination of outcome scale and predictor structure narrows the metric choice. A continuous outcome with a two-level categorical predictor points to the d family. A continuous outcome with a continuous predictor points to the r family. A categorical outcome with a categorical predictor points to the odds ratio or risk ratio family. A continuous outcome with a categorical predictor that has more than two levels points to the variance explained family.
Question 3: What is the primary research question?
The research question determines whether you need a measure of difference, a measure of association, or a measure of variance explained. A question such as "Does treatment A increase growth rate compared with control?" requires a difference measure. A question such as "Is growth rate associated with temperature?" requires an association measure. A question such as "how much of the variation in growth rate is explained by genotype?" requires a variance explained measure.
These three questions form a filter. If you cannot answer all three questions for a given outcome, you have not defined the analysis clearly enough to select an effect size. Go back to the study design and clarify the primary outcome and the comparison of interest.
The Metric Selection Matrix
The table below shows the recommended effect size metric for each combination of outcome scale and predictor structure. Use this matrix as a starting point, not as a final answer. The biological context and the conventions in your field may justify a different choice, but you must document the justification.
| Outcome Scale | Predictor Structure | Recommended Effect Size Family | Example Metric |
|---|---|---|---|
| Continuous | Two groups | d family | Cohen's d or Hedges' g |
| Continuous | More than two groups | Variance explained | Partial eta squared or omega squared |
| Continuous | Continuous | r family | Pearson's r or Spearman's rho |
| Ordinal | Two groups | r family | Rank biserial correlation |
| Ordinal | Continuous | r family | Spearman's rho |
| Categorical | Two groups | Odds ratio or risk ratio | Odds ratio with 95% confidence interval |
| Categorical | More than two groups | Variance explained or odds ratio | Cramer's V or odds ratios for pairwise comparisons |
| Categorical | Continuous | Logistic regression coefficient | Exponentiated coefficient with confidence interval |
The matrix is a starting point. For example, a study with a continuous outcome and a two group predictor can also be analyzed with a point biserial correlation, which is mathematically related to the standardized mean difference. The choice between Cohen's d and the point biserial correlation depends on how you want the reader to interpret the result. Cohen's d expresses the difference in standard deviation units, which is intuitive for a difference between two groups. The point biserial correlation expresses the strength of the association between group membership and the outcome, which is intuitive for a relationship framing.
Documenting the Decision
The decision framework produces a record that you can place in the methods section of the manuscript. The record should include four elements.
First, state the primary outcome variable and its scale. For example, "the primary outcome was the change in seedling root length after 14 days of treatment, measured in millimeters as a continuous variable."
Second, state the predictor structure. For example, "the primary predictor was the treatment group with four levels: control, low dose, medium dose, and high dose."
Third, state the effect size metric and the justification. For example, "because the outcome was continuous and the predictor had four levels, the effect size was partial eta squared, which expresses the proportion of variance in root length explained by the treatment."
Fourth, state the confidence interval level. The convention is 95%, but you should state it explicitly. For example, "all effect sizes are reported with 95% confidence intervals."
This decision record serves two purposes. It forces you to make the decision before the analysis, and it gives the reader the information needed to evaluate the choice. The EQUATOR Network reporting guidelines for your study design will indicate whether the decision record should appear in the methods section or in a supplementary file.
A Worked Example of the Decision Framework
Consider a study that measures the effect of a new fertilizer on the yield of tomato plants. The outcome is the total fruit mass per plant in grams, which is a continuous variable. The predictor is the fertilizer treatment with two levels, the new fertilizer and the standard fertilizer. The primary research question is whether the new fertilizer increases yield compared to the standard fertilizer.
The decision framework produces the following record. The outcome is continuous. The predictor is categorical with two levels. The primary question is a difference question. The effect size metric is the standardized mean difference, Cohen's d. The confidence interval is 95 percent.
The methods section would state: "The primary outcome was total fruit mass per plant in grams. The primary predictor was the fertilizer treatment with two levels. The effect size metric was Cohen's d, calculated as the difference between the group means divided by the pooled standard deviation. The effect size is reported with a 95% confidence interval."
Now consider a different study. The outcome variable is the presence or absence of a fungal infection on each plant, which is a categorical variable. The predictor is the same fertilizer treatment with two levels. The primary question is whether the new fertilizer changes the odds of infection.
The decision record changes. The outcome is categorical. The predictor is categorical with two levels. The metric is the odds ratio. The methods section would state: "The primary outcome is the presence of fungal infection, a binary variable. The primary predictor is the fertilizer treatment with two levels. The effect size metric is the odds ratio, reported with a 95% confidence interval."
The two examples show how the same study with a different outcome variable produces a different effect size metric. The decision framework forces the researcher to match the metric to the outcome and the question.
Verification Method for the Selected Metric
After you calculate the effect size, verify that the metric is appropriate for the data structure. The verification method has three checks.
Check 1: Confirm the metric is bounded correctly.
The correlation coefficient is bounded between -1 and 1. The variance explained measures are bounded between 0 and 1. The odds ratio is bounded between 0 and infinity. If your calculated effect size falls outside the expected range, the calculation is wrong or the metric is not appropriate for the data.
Check 2: Confirm the metric is consistent with the p value direction.
The effect size and the p value must tell the same story. If the p value is below 0.05 and the effect size is near zero, the result is inconsistent. If the p value is above 0.05 and the effect size is large, the result is also inconsistent. The inconsistency indicates a calculation error or a mismatch between the effect size and the statistical test.
Check 3: Confirm the metric is consistent with the biological direction.
The sign of the effect size must match the observed direction of the effect. If the treatment group has a higher mean than the control group, the standardized mean difference must be positive. If the treatment group has a lower mean, the standardized mean difference must be negative. A mismatch indicates a coding error in the analysis.
Check 4: Confirm the confidence interval is plausible.
The confidence interval must contain the point estimate. The interval should not include impossible values. For example, a correlation coefficient confidence interval should not include values below -1 or above 1. An odds ratio confidence interval should not include negative values.
The verification method is a quality control step. It does not require a second researcher or a second software package. It is a logical check that you can perform on your own results. The National Library of Medicine Research Methods Resources provides access to statistical texts that describe the properties of each effect size metric, which you can use to confirm the expected range and behavior of the metric.
Common Failure Patterns in Metric Selection
The decision framework addresses several common failure patterns that appear in biological manuscripts.
Failure Pattern 1: Using the default metric from the software.
Statistical software often provides a default effect size for each analysis. The default may not be the most appropriate metric for the research question. For example, a software package may output partial eta squared for an analysis of variance, but the researcher may need a standardized mean difference for a specific comparison. The decision framework requires the researcher to select the metric before the analysis, not to accept the software default.
Failure Pattern 2: Using a difference metric for an association question.
A researcher who asks whether two variables are associated but reports a standardized mean difference has selected the wrong metric. The standardized mean difference describes the difference between groups, not the association between variables. The decision framework prevents this error by requiring the researcher to state the primary research question before selecting the metric.
Failure Pattern 3: Using an association metric for a difference question.
A researcher who asks whether two groups differ but reports a correlation coefficient has selected the wrong metric. The correlation coefficient describes the association between variables, not the difference between groups. The decision framework prevents this error by requiring the researcher to state the primary research question.
Failure Pattern 4: Changing the metric after the analysis.
A researcher who changes the effect size metric after seeing the results is at risk of selecting the metric that produces the most favorable result. This practice is a form of selective reporting. The decision framework prevents this error by requiring the researcher to document the metric before the analysis. The Committee on Publication Ethics Core Practices describe the expectations for responsible reporting, which includes avoiding selective reporting.
Failure Pattern 5: Not documenting the metric in the methods section.
A manuscript that reports an effect size in the results section but does not state the metric in the methods section leaves the reader unable to interpret the result. The decision framework prevents this error by requiring the researcher to write the decision record in the methods section.
When to Escalate to a Biostatistician
The decision framework covers the common study designs in biological research. Some designs require the input of a biostatistician. Escalate to a biostatistician when the study design includes any of the following features.
The study has a hierarchical or nested structure, such as multiple measurements per individual or individuals nested within groups. The effect size for a mixed effects model is not the same as the effect size for a simple two group comparison.
The study has a longitudinal structure with repeated measurements over time. The effect size for a repeated measures analysis requires a different calculation than the effect size for a single time point.
The study has a multivariate outcome with multiple correlated dependent variables. The effect size for a multivariate analysis is not the same as the effect size for a univariate analysis.
The study has a survival outcome with censored data. The effect size for a survival analysis is the hazard ratio, which is not the same as the odds ratio or the risk ratio.
The study has a complex sampling design with weighting or clustering. The effect size for a complex sample design requires the design to be accounted for in the calculation.
The biostatistician can provide guidance on the appropriate effect size metric and the calculation method. The consultation should be documented in the study protocol and the methods section of the manuscript.
The Decision Record as a Reproducibility Tool
The decision record serves a second purpose beyond the methods section. It is a reproducibility tool. A reader who wants to reproduce the analysis needs to know the effect size metric, the calculation method, and the software version. The decision record provides this information.
The NIH Data Management and Sharing Policy describes the expectations for data sharing in NIH funded research. The policy requires that data be shared in a way that is consistent with the principles of reproducibility. The decision record is part of the analysis documentation that supports reproducibility.
The decision record also supports the inclusion of the study in future meta-analyses. A meta-analysis combines effect sizes from multiple studies. The meta-analysis can only include studies that report the effect size in a comparable metric. The decision record ensures that the metric is stated clearly, which allows the meta-analyst to determine whether the study can be included.
The Decision Framework in the Context of Reporting Guidelines
The EQUATOR Network maintains a collection of reporting guidelines for different study designs. The guidelines for each study design include specific requirements for effect size reporting. The decision framework is a tool that helps the researcher meet the requirements of the guideline.
The decision framework does not replace the reporting guideline. The reporting guideline specifies what to report. The decision framework specifies how to select the metric. The researcher should use both the reporting guideline and the decision framework.
The decision framework is also consistent with the expectations of the NIH Grants and Funding resources. The grant review process evaluates whether the proposed study design is rigorous. The decision framework demonstrates that the researcher has considered the effect size metric before the data collection, which is a sign of rigorous design.
The Decision Framework in the Laboratory Notebook
The decision record should be written in the laboratory notebook at the same time as the study protocol. The notebook entry should include the following information.
The date of the decision. The outcome variable and its scale. The predictor variable and its structure. The primary research question. The selected effect size metric. The justification for the metric. The confidence interval level. The statistical software and version.
The notebook entry provides a dated record of the decision. The record is useful if a question arises about the selection of the metric during the peer review process. The record also provides the information needed to write the methods section of the manuscript.
The Decision Framework in the Peer Review Process
The decision framework is also useful for peer reviewers. A reviewer who reads a manuscript can apply the three question filter to the reported effect size. If the reported effect size does not match the outcome scale and the predictor structure, the reviewer can request a justification.
The reviewer can also check the decision record in the methods section. If the methods section does not state the effect size metric, the reviewer can request the statement. The reviewer can also check the verification method by confirming that the effect size is within the expected range and is consistent with the p value and the direction of the data.
The decision framework is a shared tool for the researcher and the reviewer. It provides a common language for discussing the effect size selection and the reporting.
Frequently Asked Questions
What is the difference between an effect size and a p-value?
A p-value indicates the probability of observing the data if the null hypothesis is true. An effect size indicates the magnitude of the difference or the association. The p-value does not indicate the size of the effect, and the effect size does not indicate the statistical significance.
How do I choose the right effect size metric for my study?
The choice of the effect size metric depends on the study design and the type of data. For two-group comparisons with continuous outcomes, use a standardized mean difference. For correlations, use a correlation coefficient. For categorical outcomes, use an odds ratio. For analysis of variance designs, use partial eta squared or omega squared.
What is a good effect size?
There is no universal answer to this question. The interpretation of the effect size depends on the biological context and the study field. A small effect size in one context may be a large effect size in another context. The effect size should be interpreted in the context of the study.
How do I calculate the confidence interval for an effect size?
The confidence interval for an effect size can be calculated using the statistical software or using the formulas for the specific effect size metric. The confidence interval provides a range of plausible values for the effect size.
Should I report effect sizes for secondary analyses?
Yes, the effect sizes for secondary analyses should be reported when they support the primary findings. The effect size for a secondary analysis can be important for the interpretation of the primary findings.
How do I report effect sizes in a figure?
The effect size can be reported in a figure by including the effect size estimate and the confidence interval in the figure legend or in the figure itself. The effect size can also be displayed in a forest plot or a similar graphical format.
What should I do if the effect size is not available in my statistical software?
If the effect size is not available in the statistical software, the effect size can be calculated using the formulas for the effect size metric. The formulas are available in statistical textbooks and in the Research Methods Resources from the National Library of Medicine.
How do I interpret an odds ratio as an effect size?
An odds ratio indicates the odds of an outcome in one group relative to the odds in another group. An odds ratio of 1 indicates no association. An odds ratio greater than 1 indicates increased odds, and an odds ratio less than 1 indicates decreased odds. The magnitude of the odds ratio should be interpreted in the context of the baseline risk of the outcome.
Related Bioinformatics Guides
- RNA-Seq Batch Effect Detection and Correction
- Data Stewardship vs. Data Governance: Roles and Responsibilities in Research
- Research Data Stewardship: Benefits and Implementation Strategies
- Metabolomics Data Analysis Workflow: From Raw Data to Biological Insight
- Metagenome Assembled Genome Analysis: From Bins to Biological Insights
Related Clinical & Scientific Guides
- A Practical Guide to Detecting Antimicrobial Resistance Genes in Shotgun Metagenomic Data
- Computational Immunology: Modeling the Immune System
- How to Set Hard Filters for Germline Variant Calling: A Practical Guide to GATK Best Practices
References and Further Reading
- Research Methods Resources. National Library of Medicine.
- EQUATOR Network. EQUATOR Network.
- Core Practices. Committee on Publication Ethics.
- NIH Grants and Funding. National Institutes of Health.
- ORCID for Researchers. ORCID.
- Data Management and Sharing Policy. National Institutes of Health.
- NCBI Data Resources. National Center for Biotechnology Information.
- EMBL-EBI Training. European Bioinformatics Institute.
- Sample size determination and power analysis using the G*Power software.. Journal of educational evaluation for health professions, 2021.
- A Simple Guide to Effect Size Measures.. JAMA otolaryngology-- head & neck surgery, 2023.
- Progressive statistics for studies in sports medicine and exercise science.. Medicine and science in sports and exercise, 2009.
This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.