Effect Size in Statistics: Why It Matters and How to Interpret It
A statistically significant result tells you that an effect probably exists, but it does not tell you how large that effect is. Effect size quantifies the magnitude of a difference or the strength of an association between variables, and it is essential for determining practical or theoretical importance, calculating statistical power, and combining results across studies in meta-analysis. This article defines effect size, contrasts it with p-values, and provides interpretation guidelines for common metrics including Cohen's d, eta squared, and odds ratios, with a practical example of calculating and reporting effect size.
At a Glance
Effect size answers a different question than statistical significance. A p-value indicates whether an observed effect is likely due to chance, while an effect size indicates how meaningful that effect is in practical terms. Researchers and practitioners should report both, along with confidence intervals, to provide a complete picture of their findings.
| Effect Size Measure | Family | What It Quantifies | Common Interpretation Thresholds | Typical Use Case |
|---|---|---|---|---|
| Cohen's d | d family | Standardized mean difference between two groups | 0.20 small, 0.50 medium, 0.80 large | Comparing treatment versus control group means |
| Eta squared (η²) | r family | Proportion of variance explained by a factor | 0.01 small, 0.06 medium, 0.14 large | Analysis of variance (ANOVA) designs |
| Odds ratio | d family | Strength of association in binary outcomes | 1.0 no effect, above 1.0 increased odds, below 1.0 decreased odds | Case-control studies, logistic regression |
The choice of effect size measure depends on the research question, study design, target audience, and statistical assumptions. Nonstandardized effect sizes, such as mean differences, are more informative and easier to interpret in the context of clinical or practical relevance. Standardized effect sizes are unitless and useful for comparing effects across different outcome measures or studies.
What Effect Size Means in Statistical Analysis
Effect size is a quantitative reflection of the magnitude of some phenomenon that is used for the purpose of addressing a question of interest. This definition is intentionally broad and encompasses many specific measures. The term has three facets: dimension, measure or index, and value. The dimension refers to what is being measured, the measure or index is the specific statistic used, and the value is the numerical result.
Effect sizes can be organized into two families. The d family includes measures that quantify differences between groups, such as Cohen's d and Hedges' g. The r family includes measures that quantify the strength of association, such as correlation coefficients and eta squared. Understanding which family a measure belongs to helps researchers select the appropriate statistic for their design.
Nonstandardized effect sizes are presented in the same units as the characteristic being measured. For example, a mean difference in blood pressure of 5 mmHg is a nonstandardized effect size. These have the advantage of being more informative and easier to evaluate in light of clinical significance or practical relevance. Standardized effect sizes are unitless and helpful for combining and comparing effects of different outcome measures or across different studies in meta-analysis.
Why Effect Size Matters More Than P-Values Alone
Null hypothesis significance testing is the dominant statistical approach in biology, yet it has many frequently unappreciated problems. Most importantly, this approach does not provide two crucial pieces of information: the magnitude of an effect of interest and the precision of the estimate of that magnitude. Researchers should ultimately be interested in biological importance, which may be assessed using the magnitude of an effect but not its statistical significance.
Drawing conclusions based on a dichotomous p-value instead of a spectrum can mislead researchers into concluding there is no difference between two groups or two treatments when a meaningful difference actually exists. The utilization of effect size, or the magnitude of difference between studied groups, helps obtain a better global understanding of the statement no effect. Although statistical significance does not mean clinical significance, adequate interpretation of data discloses transparent results and conclusions.
Combined use of an effect size and its confidence interval enables assessment of relationships within data more effectively than the use of p-values, regardless of statistical significance. Routine presentation of effect sizes encourages researchers to view their results in the context of previous research and facilitates incorporation of results into future meta-analysis, which has become a standard method of quantitative review.
Cohen's d: Standardized Mean Difference
Cohen's d is the most commonly reported effect size estimate for t tests. It quantifies the standardized difference between two group means. The calculation involves dividing the difference between group means by the pooled standard deviation, producing a unitless value that can be compared across studies using different measurement scales.
Interpreting Cohen's d Values
Researchers typically use Cohen's guidelines of d = 0.20, 0.50, and 0.80 to interpret observed effect sizes as small, medium, or large, respectively. However, these guidelines were not based on quantitative estimates and are only recommended if field-specific estimates are unknown. Field-specific guidelines may differ substantially from these general benchmarks.
For example, a study in gerontology found that Cohen's guidelines appear to overestimate effect sizes in that field. Researchers in gerontology are encouraged to use Cohen's d or Hedges' g values of 0.15, 0.40, and 0.75 to interpret small, medium, and large effects, and to recruit larger samples accordingly. This illustrates why researchers should seek field-specific benchmarks when available.
Practical Example of Cohen's d
Consider a study comparing two treatments for pediatric phimosis. The study reported that balloon catheter dilation treatment had a mean operative time of 5.03 minutes compared to 29.38 minutes for conventional circumcision, with a Cohen's d of -5.13. This very large effect size indicates that the difference in operative time between the two procedures is substantial and practically meaningful, far exceeding the conventional large threshold of 0.80.
Similarly, healing time showed a Cohen's d of -3.47, indicating that the balloon catheter approach produced much faster healing. These effect sizes provide information that p-values alone cannot convey. While the p-values indicated statistical significance, the effect sizes revealed the magnitude of the practical advantage.
Eta Squared and Partial Eta Squared for ANOVA
Eta squared is the most commonly reported effect size estimate for analysis of variance. It represents the proportion of total variance in the dependent variable that is attributable to the independent variable or factor. Partial eta squared is a variant that accounts for other factors in the model and is frequently reported in factorial designs.
Interpreting Eta Squared Values
Common interpretation thresholds for eta squared are 0.01 for a small effect, 0.06 for a medium effect, and 0.14 for a large effect. These values represent the proportion of variance explained, so an eta squared of 0.14 indicates that 14 percent of the variance in the outcome is explained by the factor of interest.
A study of undergraduate nursing students comparing a BOPPPS model combined with case-based learning against traditional lecture-based learning reported a partial eta squared of 0.669 for post-test knowledge scores. This indicates that approximately 67 percent of the variance in post-test scores was explained by the teaching method after adjusting for pre-test scores, a very large effect.
Eta Squared in Observational Studies
Eta squared is also useful in observational research. A study of public information-seeking behavior in primary care used eta squared to evaluate seasonal variation in search interest for family medicine terms. The analysis found significant seasonal variation for the term family physician with an eta squared of 0.232, indicating that season explained about 23 percent of the variance in search interest.
Another study of dental student placements used eta squared to assess differences across regions. The analysis found an eta squared of 0.21 for overall student placement consistency across regions, indicating a medium to large effect. This type of information helps administrators understand whether regional differences are practically meaningful or merely statistically detectable.
Odds Ratios and Other Association Measures
Odds ratios quantify the strength of association between variables in studies with binary outcomes. An odds ratio of 1.0 indicates no association, values above 1.0 indicate increased odds of the outcome, and values below 1.0 indicate decreased odds. Odds ratios are commonly reported in case-control studies and logistic regression analyses.
The interpretation of odds ratios requires attention to the direction and magnitude of the effect. An odds ratio of 2.0 indicates that the odds of the outcome are twice as high in the exposed group compared to the unexposed group. The clinical or practical significance of an odds ratio depends on the baseline risk and the context of the decision being made.
Confidence intervals are essential companions to odds ratios and other effect sizes. A wide confidence interval indicates substantial uncertainty about the true effect size, while a narrow interval indicates more precise estimation. When a confidence interval includes the null value, such as 1.0 for an odds ratio or 0 for a mean difference, the evidence is compatible with no effect.
Confidence Intervals and Precision of Effect Size Estimates
Effect sizes are estimates of population parameters, and they carry uncertainty. Reporting a confidence interval acknowledges the uncertainty with which the population value of the effect size has been estimated. For a complete and meaningful interpretation of results, investigators should make clear the type of effect size being reported, its magnitude and direction, the degree of uncertainty as presented by the confidence intervals, and whether the results are compatible with a clinically meaningful effect.
The Publication Manual of the American Psychological Association calls for the reporting of effect sizes and their confidence intervals. Despite this recommendation, a survey of articles published in 2009 and 2010 in the Journal of Experimental Psychology: General found that effect sizes were reported for fewer than half of the analyses, and no article reported a confidence interval for an effect size. This reporting gap persists despite professional guidelines.
Confidence intervals serve multiple purposes. They indicate the precision of the effect size estimate, they allow readers to assess whether the results are compatible with meaningful effects, and they facilitate meta-analysis by providing information about the variability of estimates across studies.
Practical Workflow for Calculating and Reporting Effect Size
Researchers should follow a systematic process for incorporating effect sizes into their statistical practice. This workflow applies to experimental studies, observational research, and clinical investigations.
Step 1: Select the Appropriate Effect Size Measure
The choice of effect size measure depends on the research question, study design, target audience, and statistical assumptions. For comparing two group means, Cohen's d is appropriate. For analysis of variance designs, eta squared or partial eta squared is standard. For binary outcomes, odds ratios or risk differences are suitable. For correlation-based designs, Pearson's r is the natural choice.
Step 2: Calculate the Effect Size and Confidence Interval
Calculate the effect size using the appropriate formula for the selected measure. For Cohen's d, divide the difference between group means by the pooled standard deviation. For eta squared, divide the sum of squares for the factor by the total sum of squares. For odds ratios, exponentiate the logistic regression coefficient. Then calculate the confidence interval using the appropriate standard error.
Step 3: Interpret the Effect Size in Context
Interpret the effect size using field-specific benchmarks when available, or general guidelines when field-specific values are unknown. Consider whether the magnitude of the effect is practically or clinically meaningful, also statistically significant. For nonstandardized effect sizes, evaluate the effect in the context of the measurement units and the decisions that depend on the results.
Step 4: Report the Effect Size With the P-Value
Report the effect size alongside the p-value and confidence interval in all results sections. Include the type of effect size, its magnitude and direction, and the confidence interval. This practice enables readers to assess practical significance and facilitates meta-analysis.
Step 5: Use Effect Sizes for Power Analysis
Effect sizes are needed for sample size calculation and power analysis. A priori power analyses require an assumed effect size to determine the sample size needed to detect an effect with adequate power. Researchers can use effect sizes from previous studies, pilot data, or field-specific benchmarks for these calculations.
Records and Measurements for Effect Size Reporting
Accurate effect size reporting depends on complete and accurate data records. Researchers should document the following information for each analysis:
| Record Element | Description | Purpose |
|---|---|---|
| Group means and standard deviations | Summary statistics for each group | Calculate Cohen's d and other mean-based effect sizes |
| Sample sizes per group | Number of observations in each group | Calculate pooled standard deviations and confidence intervals |
| Sum of squares values | ANOVA summary table values | Calculate eta squared and partial eta squared |
| Regression coefficients and standard errors | Model output from regression analyses | Calculate odds ratios and standardized coefficients |
| Confidence intervals for effect sizes | Interval estimates for each effect size | Report precision and facilitate meta-analysis |
Maintaining these records ensures that effect sizes can be calculated consistently and that confidence intervals can be derived. Documentation also supports reproducibility and enables other researchers to verify calculations.
Common Failure Patterns in Effect Size Reporting
Several common problems undermine the quality of effect size reporting in research. Recognizing these patterns helps researchers avoid them and helps readers identify weaknesses in published studies.
Reporting Only Statistical Significance
The most common failure is reporting p-values without effect sizes. A survey of articles in the Journal of Experimental Psychology: General found that effect sizes were reported for fewer than half of the analyses. For t tests, two-thirds of the articles did not report an associated effect size estimate. This pattern deprives readers of information about the magnitude of effects.
Using Inappropriate Interpretation Thresholds
Applying Cohen's general guidelines without considering field-specific benchmarks can lead to misinterpretation. Research in gerontology found that Cohen's guidelines overestimate effect sizes in that field, and researchers are encouraged to use field-specific values. Researchers should seek field-specific benchmarks when available instead of defaulting to general guidelines.
Confusing Statistical and Practical Significance
Statistical significance does not mean clinical significance. A small effect can be statistically significant with a large sample, while a large effect may fail to reach significance with a small sample. Researchers must interpret effect sizes in the context of practical or clinical relevance, not rely solely on p-values.
Omitting Confidence Intervals
Many researchers report effect sizes without confidence intervals. The survey of psychology articles found that no article reported a confidence interval for an effect size. This omission prevents readers from assessing the precision of the estimate and the compatibility of the results with meaningful effects.
Mislabeling Effect Size Measures
Confusion in the literature about the definition of effect size leads to inconsistent use of the term. Researchers should clearly identify the specific measure being reported, such as Cohen's d or partial eta squared, instead of using the generic term effect size without specification.
Limitations and Caveats in Effect Size Interpretation
Effect sizes are valuable tools, but they have limitations that researchers must acknowledge. Understanding these limitations prevents overinterpretation and supports appropriate use.
Context Dependence of Interpretation Thresholds
General interpretation guidelines such as Cohen's thresholds are not universal. Field-specific estimates can differ substantially from general benchmarks. Researchers should use field-specific values when available and treat general guidelines as rough starting points instead of definitive rules.
Sensitivity to Study Design
Effect sizes are influenced by study design characteristics. For example, the magnitude of an effect size can depend on the variability within groups, the measurement instrument used, and the specific comparison being made. Researchers should consider whether the effect size is comparable across studies with different designs.
Nonstandardized Versus Standardized Measures
Nonstandardized effect sizes are more informative for practical decisions but cannot be compared across different measurement scales. Standardized effect sizes enable comparison across studies but may obscure the practical meaning of the effect. Researchers should report both when possible.
Uncertainty in Effect Size Estimates
Effect sizes estimated from samples carry sampling error. Confidence intervals communicate this uncertainty, but many studies report effect sizes without intervals. Readers should be cautious when interpreting effect sizes without accompanying confidence intervals.
Quality Controls and Professional Escalation Criteria
Researchers should implement quality controls to ensure accurate effect size calculation and reporting. When discrepancies or concerns arise, escalation to appropriate professionals is warranted.
Verification of Calculations
Independent verification of effect size calculations helps prevent errors. Researchers should check their calculations against published examples or use validated statistical software. When effect sizes seem implausibly large or small, recalculate and verify the underlying data.
Consultation With Statistical Experts
When selecting effect size measures for complex designs or when interpreting results with unusual characteristics, consultation with a statistician or methodological expert is appropriate. This is particularly important for meta-analyses, where consistent effect size calculation across studies is essential.
Adherence to Reporting Guidelines
Researchers should follow established reporting guidelines for their study type. The EQUATOR Network provides access to reporting guidelines for various study designs, and resources such as the Experimental Design Assistant from NC3Rs support rigorous study planning. Following these guidelines helps ensure that effect sizes and other statistical information are reported completely.
Escalation for Data Quality Issues
When data quality issues affect effect size calculations, such as missing data, outliers, or measurement errors, escalate the issue to the appropriate data manager or statistical consultant. Do not proceed with effect size reporting until data quality concerns are resolved.
Safety and Regulatory Context for Effect Size Reporting
Effect size reporting has implications for research integrity and regulatory compliance. Funding agencies, journals, and regulatory bodies increasingly require transparent statistical reporting.
Journal Requirements
Many journals now require effect size reporting as a condition of publication. The Publication Manual of the American Psychological Association calls for the reporting of effect sizes and their confidence intervals. Researchers should check journal-specific requirements before submission.
Research Data Management
The National Institute of Standards and Technology supports the Research Data Framework, which provides guidance on managing research data throughout its lifecycle. Proper data management supports accurate effect size calculation and enables verification by other researchers.
Literature Search and Evidence Synthesis
Effect sizes are essential for meta-analysis, which has become a standard method of quantitative review. The National Center for Biotechnology Information provides literature resources including PubMed, which researchers use to identify studies for evidence synthesis. Complete effect size reporting in primary studies facilitates these secondary analyses.
Frequently Asked Questions
What is the difference between statistical significance and effect size?
Statistical significance indicates whether an observed effect is likely due to chance, typically assessed using a p-value. Effect size quantifies the magnitude of the difference or the strength of the association. A result can be statistically significant with a very small effect, especially with large samples, or fail to reach significance with a large effect in a small sample. Both pieces of information are needed for complete interpretation.
How do I interpret Cohen's d values?
Cohen's d values of 0.20, 0.50, and 0.80 are commonly interpreted as small, medium, and large effects, respectively. However, these guidelines were not based on quantitative estimates and are only recommended if field-specific estimates are unknown. Some fields have developed their own benchmarks, such as gerontology where values of 0.15, 0.40, and 0.75 are recommended.
What is the difference between eta squared and partial eta squared?
Eta squared represents the proportion of total variance in the dependent variable explained by the independent variable. Partial eta squared represents the proportion of variance explained by the independent variable after accounting for other factors in the model. Partial eta squared is commonly reported in factorial ANOVA designs where multiple factors are included.
When should I use an odds ratio instead of Cohen's d?
Use an odds ratio when the outcome is binary, such as presence or absence of a condition, and you want to quantify the association between an exposure or treatment and that outcome. Use Cohen's d when comparing means of a continuous outcome between two groups. The choice depends on the nature of the outcome variable and the research question.
Why should I report confidence intervals with effect sizes?
Confidence intervals communicate the precision of the effect size estimate and acknowledge the uncertainty with which the population value has been estimated. They allow readers to assess whether the results are compatible with clinically meaningful effects and facilitate meta-analysis. Reporting effect sizes without confidence intervals omits crucial information about estimate precision.
Can effect sizes be compared across different studies?
Standardized effect sizes such as Cohen's d and eta squared can be compared across studies because they are unitless. This comparability makes them useful for meta-analysis. Nonstandardized effect sizes are presented in original measurement units and cannot be directly compared across studies using different measures.
What effect size should I use for sample size calculations?
Use the effect size that is appropriate for your primary analysis and the expected magnitude of the effect you want to detect. This can be based on previous studies, pilot data, or field-specific benchmarks. A priori power analyses require an assumed effect size to determine the sample size needed for adequate power.
How do I report effect sizes in my results section?
Report the type of effect size, its magnitude and direction, and its confidence interval alongside the p-value. For example, report Cohen's d with its confidence interval for t tests and partial eta squared with its confidence interval for ANOVA. Make clear which measure you are reporting and interpret the magnitude in the context of your field.
Related Articles
- What Is Animal Imprinting? How It Works and Why It Matters
- Are Insects Animals? The Surprising Answer and Why It Matters
- Symbiosis in Nature: Types, Examples, and Why It Matters
- Wolf Size Comparison: How Big Are Wolves Really?
- Filter Feeding Animals: How They Eat and Why It Matters
References and Further Reading
- Research Data Framework. National Institute of Standards and Technology.
- EQUATOR Network. EQUATOR Network.
- Experimental Design Assistant. NC3Rs.
- NCBI Literature Resources. National Center for Biotechnology Information.
- PubMed. National Library of Medicine.
- Effect size estimates: current use, calculations, and interpretation.. Journal of experimental psychology. General, 2012.
- A Simple Guide to Effect Size Measures.. JAMA otolaryngology-- head & neck surgery, 2023.
- Effect size, confidence interval and statistical significance: a practical guide for biologists.. Biological reviews of the Cambridge Philosophical Society, 2007.
- Effect Size Guidelines, Sample Size Calculations, and Statistical Power in Gerontology.. Innovation in aging, 2019.
- The effect of display size on ultrasound interpretation.. The American journal of emergency medicine, 2022.
- Application and interpretation of linear-regression analysis.. Medical hypothesis, discovery & innovation ophthalmology journal, 2024.
- On effect size.. Psychological methods, 2012.
- Editorial Commentary: The Power of Interpretation: Utilizing the P Value as a Spectrum, in Addition to Effect Size, Will Lead to Accurate Presentation of Results.. Arthroscopy : the journal of arthroscopic & related surgery : official publication of the Arthroscopy Association of North America and the International Arthroscopy Association, 2022.
- A novel non-invasive technique for pediatric phimosis treatment: assessing efficacy and post-treatment family care - a retrospective cohort study.. 2026.
- Effects of the BOPPPS model combined with case-based learning on knowledge acquisition and learning engagement in undergraduate nursing students: a multi-cohort quasi-experimental study.. 2026.
- Public Information-Seeking Behavior in Primary Care: A National Case Study of Türkiye.. 2026.
- Agreement between patient self-reported and assessor-rated frailty using the Clinical Frailty Scale.. 2026.
- Effectiveness of bullying prevention-associated interventions among children and adolescents: an umbrella review of systematic review and meta-analysis.. 2026.
- Eta Squared and Partial Eta Squared as Measures of Effect Size in Educational Research.. 2011.
- Pion-eta scalar-isovector 3-coupled channel amplitude fitted to branching ratios and threshold plus subthreshold parameters. 2017.
- Undergraduate dental student placement in primary health care corporation in Qatar: a provider perspective. BMC Medical Education, 2026.
- Inter-daily variability in body composition among young men. Journal of Physiological Anthropology, 2015.
- The effect of heritage tourism interpretation media type on tourists’ eWOM: the moderating role of travel group size. Journal of Sustainable Tourism, 2025.
- Novel Effect Size Interpretation Guidelines and an Evaluation of Statistical Power in Rehabilitation Research. Archives of Physical Medicine and Rehabilitation, 2020.
- A test of two interpretations of the apparent size effects in a distorted room. Journal of Experimental Psychology, 1962.
This article is educational and does not replace institutional policy, professional advice, or applicable safety and regulatory requirements.