Cohen's d vs hedges' g

By Dr. Zubair Khalid, DVM, MS, PhD ·

Cohen's d vs hedges' g

Key Takeaways

  • Cohen's d and Hedges' g both quantify standardized mean differences, but Hedges' g incorporates a bias correction factor crucial for small sample sizes prevalent in biological research, where sample standard deviations tend to underestimate population standard deviations.
  • For studies with fewer than 20 subjects per group, Hedges' g is recommended to prevent overestimation of the true effect size, as the bias in Cohen's d is most pronounced in these low-N scenarios.
  • When sample sizes exceed 50 subjects per group, the bias correction in Hedges' g becomes negligible, and Cohen's d provides a sufficiently accurate estimate, often aligning closely with Hedges' g.
  • The choice between Cohen's d and Hedges' g should be predetermined and documented in the study protocol or preregistration before data collection to ensure methodological rigor and prevent post-hoc bias.
  • Hedges' g is particularly advantageous for confirmatory studies, pilot data, and meta-analyses involving studies with varying sample sizes, as it offers a more accurate and consistent effect size estimate across diverse biological experimental designs.
  • The magnitude of the Hedges' correction is inversely related to the total degrees of freedom (sum of sample sizes minus two), meaning smaller total sample sizes result in a more substantial correction factor, reducing the effect size estimate compared to Cohen's d.

Quick Answer

  • Cohen's d and Hedges' g both measure standardized mean differences, but Hedges' g applies a correction factor that reduces bias in small samples common in biological research.
  • For studies with fewer than 20 subjects per group, use Hedges' g to avoid overestimating the true effect size.
  • Both metrics remain valid for larger samples, yet the choice should be documented in your analysis plan before data collection begins.

At a Glance

MetricBias CorrectionRecommended Sample SizePrimary Use Case
Cohen's dNoneLarge samples (n > 50 per group)Exploratory analyses, large datasets, meta-analyses with uniform metrics
Hedges' gYes, applies correction factorSmall samples (n < 20 per group)Confirmatory studies, pilot data, underpowered designs common in biology
BothDepends on computationAny sample sizeWhen reporting standardized effect sizes alongside confidence intervals

Understanding Standardized Effect Sizes in Biological Research

Standardized effect sizes allow researchers to express the magnitude of a difference between two groups in units that are comparable across studies. In biological research, where measurements often come from different scales, instruments, or laboratories, the raw difference between group means is rarely interpretable on its own. A standardized effect size converts that raw difference into a unitless quantity, typically by dividing the observed difference by a measure of variability within the groups.

The two most frequently encountered standardized effect sizes for comparing two group means are Cohen's d and Hedges' g. Both metrics answer the same question: how large is the difference between two groups relative to the variability within those groups? The distinction lies in how each metric handles the estimation of the population standard deviation from sample data.

Cohen's d uses the pooled standard deviation of the two groups as the denominator. This pooled estimate is calculated from the sample variances, and it assumes that the two groups share a common population variance. When sample sizes are large, the sample standard deviation closely approximates the population standard deviation, and Cohen's d provides an accurate estimate of the true effect size.

Hedges' g applies a correction factor to Cohen's d. This correction, often called the Hedges correction, adjusts for the fact that sample standard deviations tend to underestimate population standard deviations, particularly when sample sizes are small. The correction factor is always less than one, meaning Hedges' g is always slightly smaller than Cohen's d computed from the same data. The magnitude of the correction depends on the degrees of freedom, which are determined by the total sample size across both groups.

For biological researchers, the practical consequence of this distinction is straightforward. Many biological experiments use small sample sizes due to cost, ethical constraints, or the availability of biological material. In these situations, Cohen's d systematically overestimates the true effect size, and Hedges' g provides a more accurate estimate.

Why Small Sample Sizes Create Bias in Effect Size Estimates

The bias in Cohen's d arises from the mathematics of standard deviation estimation. The sample standard deviation is calculated using n minus 1 in the denominator, which corrects for the fact that the sample mean is itself an estimate. However, this correction does not fully eliminate the bias in the standard deviation estimate when the sample size is small. The sample standard deviation tends to underestimate the population standard deviation, and this underestimation becomes more pronounced as the sample size decreases.

When the sample standard deviation is underestimated, the denominator of Cohen's d is too small, and the resulting effect size is too large. This means that Cohen's d overestimates the true effect size in small samples. The overestimation is not constant across sample sizes. It is most severe when the total sample size is less than 20, and it diminishes as the sample size approaches 50 or more.

Hedges' g addresses this problem by multiplying Cohen's d by a correction factor that depends on the degrees of freedom. The correction factor is derived from the gamma function and is approximately equal to one minus three divided by four times the degrees of freedom minus one. For practical purposes, the correction factor is close to one when the degrees of freedom are large, and it becomes noticeably smaller than one when the degrees of freedom are small.

The practical implication for biological research is that the choice between Cohen's d and Hedges' g should be made with the sample size in mind. A researcher who reports Cohen's d from a study with eight animals per group is likely to report an effect size that is larger than the true effect. A researcher who reports Hedges' g from the same data will report a more accurate estimate.

The Mathematical Relationship Between Cohen's d and Hedges' g

Cohen's d is calculated as the difference between the two group means divided by the pooled standard deviation. The pooled standard deviation is the square root of the weighted average of the two group variances, where the weights are the degrees of freedom from each group.

Hedges' g is calculated by applying a correction factor to Cohen's d. The correction factor is a function of the total degrees of freedom, which is the sum of the two group sample sizes minus two. The correction factor is always less than one, so Hedges' g is always smaller than Cohen's d for the same data.

The relationship between the two metrics can be expressed as Hedges' g equals Cohen's d multiplied by the correction factor. The correction factor approaches one as the total sample size increases, which means that the two metrics converge for large samples. For small samples, the correction factor is meaningfully less than one, and the difference between the two metrics is substantial.

The practical implication of this relationship is that a researcher who computes Cohen's d from a small sample and then compares that value to published Hedges' g values from other studies will be comparing values that are not directly comparable. The Cohen's d value will be systematically larger, and the difference will be more pronounced for smaller samples.

When to Use Cohen's d

Cohen's d is the appropriate choice when the sample size is large enough that the bias in the standard deviation estimate is negligible. In practice, this means studies with more than 50 observations per group. At this sample size, the correction factor for Hedges' g is so close to one that the two metrics produce nearly identical values, and the choice between them has no practical consequence.

Cohen's d is also appropriate when the researcher is working within a field or a literature that has established conventions for reporting effect sizes. Some research communities have historically reported Cohen's d, and using the same metric allows for direct comparison with published values. In these cases, the researcher should still compute Hedges' g as a sensitivity check to confirm that the bias is negligible at the given sample size.

Cohen's d is the default choice in many statistical software packages, and it is the metric that appears in the output of common statistical tests. Researchers who are using these packages should be aware that the default output is Cohen's d, and they should consider whether the sample size justifies the use of this metric or whether they should apply the Hedges correction.

When to Use Hedges' g

Hedges' g is the appropriate choice when the sample size is small, which is common in biological research. Studies with fewer than 20 observations per group are particularly susceptible to the bias in Cohen's d, and the correction factor in Hedges' g provides a meaningful adjustment.

Hedges' g is also appropriate when the researcher is planning to include the study in a meta-analysis. Meta-analyses combine effect sizes from multiple studies, and the accuracy of the combined estimate depends on the accuracy of the individual effect sizes. If some studies report Cohen's d and others report Hedges' g, the meta-analyst must convert between the two metrics, which introduces additional uncertainty. Reporting Hedges' g from the outset avoids this conversion step.

Hedges' g is the recommended choice for studies with unequal sample sizes between groups. The bias in Cohen's d is more pronounced when the group sizes are unequal, and the Hedges' correction accounts for the total degrees of freedom, which reflects the unequal group sizes.

Sample Size and the Magnitude of the Correction

The magnitude of the Hedges' correction depends on the total degrees of freedom, which is the sum of the two group sample sizes minus two. The correction factor is approximately one minus three divided by four times the degrees of freedom minus one. This formula produces a correction factor that is close to one for large degrees of freedom and noticeably smaller than one for small degrees of freedom.

For a study with 10 observations per group, the total degrees of freedom is 18, and the correction factor is approximately 0.96. This means that Hedges' g is about 4 percent smaller than Cohen's d. For a study with 5 observations per group, the total degrees of freedom is 8, and the correction factor is approximately 0.90. This means that Hedges' g is about 10 percent smaller than Cohen's d.

For a study with 30 observations per group, the total degrees of freedom is 58, and the correction factor is approximately 0.99. The difference between Cohen's d and Hedges' g is less than 1 percent, and the choice between the two metrics has no practical consequence.

These numbers illustrate the practical rule of thumb: use Hedges' g when the total sample size is less than 50, and use either metric when the total sample size is greater than 50. The exact threshold depends on the precision required by the research question, but the general principle is that the correction matters when the sample size is small.

Practical Workflow for Choosing and Computing the Correct Metric

The decision between Cohen's d and Hedges' g should be made before data collection begins, and it should be documented in the study protocol or preregistration. The following workflow provides a structured approach to this decision.

First, determine the expected sample size for each group. This determination should be based on the study design, the availability of biological material, and the ethical constraints on the number of animals or subjects. The expected sample size is the primary input for the decision between Cohen's d and Hedges' g.

Second, calculate the total degrees of freedom, which is the sum of the two group sample sizes minus two. If the total degrees of freedom is less than 50, plan to use Hedges' g. If the total degrees of freedom is greater than 50, either metric is acceptable, and the choice can be based on the conventions in the field.

Third, compute the effect size using the chosen metric. If the researcher is using statistical software, the software may provide Cohen's d by default. In this case, the researcher should apply the Hedges' correction manually or use a software option that provides Hedges' g.

Fourth, report the effect size with a confidence interval. The confidence interval provides information about the precision of the effect size estimate, and it is essential for interpreting the results. A wide confidence interval indicates that the effect size is not precisely estimated, which is common in small samples.

Fifth, document the choice of metric in the methods section of the manuscript. The documentation should include the sample sizes, the degrees of freedom, and the formula used to compute the effect size. This documentation allows other researchers to reproduce the analysis and to compare the results with other studies.

Options and Tradeoffs in Effect Size Reporting

The choice between Cohen's d and Hedges' g is one of several decisions that researchers must make when reporting effect sizes. Other decisions include whether to report the effect size with a confidence interval, whether to report the raw difference between group means, and whether to report the effect size in addition to or instead of the p-value.

Reporting the effect size with a confidence interval is strongly recommended. The confidence interval provides information about the precision of the estimate, and it allows the reader to assess whether the effect size is consistent with a meaningful biological effect. A confidence interval that includes zero indicates that the effect is not statistically significant, while a confidence interval that excludes zero indicates that the effect is statistically significant.

Reporting the raw difference between group means is also useful, particularly when the outcome measure has a natural unit that is meaningful to the reader. The raw difference provides information about the magnitude of the effect in the original units, while the standardized effect size provides information about the magnitude relative to the variability.

The tradeoff between Cohen's d and Hedges' g is a tradeoff between simplicity and accuracy. Cohen's d is simpler to compute and to explain, but it is biased in small samples. Hedges' g is slightly more complex, but it provides a more accurate estimate in the sample sizes that are common in biological research.

Observations and Measurements for Effect Size Computation

The computation of Cohen's d and Hedges' g requires the following data from each group: the sample size, the sample mean, and the sample standard deviation. These values are typically reported in the results section of a study, and they are the inputs for the effect size formulas.

The sample size is the number of observations in each group. The sample mean is the arithmetic average of the observations in each group. The sample standard deviation is the square root of the sample variance, which is calculated using n minus 1 in the denominator.

The pooled standard deviation is calculated from the individual group standard deviations and sample sizes. The pooled standard deviation is the square root of the weighted average of the individual group variances, where the weights are the degrees of freedom from each group.

The degrees of freedom for the pooled standard deviation is the sum of the two group sample sizes minus two. This value is used in the Hedges' correction factor.

The researcher should verify that the data meet the assumptions of the effect size calculation. The primary assumption is that the two groups have a common population variance. If the variances are substantially different between the groups, the pooled standard deviation may not be an appropriate measure of variability, and the researcher should consider using a different effect size metric.

Records and Documentation for Effect Size Reporting

The choice between Cohen's d and Hedges' g should be documented in the study plan, the analysis code, and the final report. This documentation serves several purposes. It allows the researcher to reproduce the analysis, it allows the reader to understand the analysis, and it allows the meta-analyst to combine the results with other studies.

The study plan should state the expected sample size and the planned effect size metric. This statement should be made before data collection begins, and it should be included in the preregistration plan if the study is preregistered.

The analysis code should include the formula used to compute the effect size. If the researcher is using statistical software, the code should specify whether the output is Cohen's d or Hedges' g. If the researcher is computing the effect size manually, the code should include the correction factor.

The results section should report the effect size with a confidence interval, and it should state the metric that was used. The results section should also report the sample sizes, the means, and the standard deviations for each group, so that the reader can verify the effect size calculation.

Quality Controls for Effect Size Computation

The computation of Cohen's d and Hedges' g is straightforward, but errors can occur in the calculation. The following quality checks can help to identify and correct errors.

First, verify the sample sizes for each group. The sample sizes should match the number of observations in the data set, and they should be consistent with the degrees of freedom reported in the analysis.

Second, verify the means and standard deviations for each group. These values should be consistent with the raw data, and they should be checked against the output of the statistical software.

Third, verify the pooled standard deviation. The pooled standard deviation should be between the individual group standard deviations, and it should be closer to the standard deviation of the larger group.

Fourth, verify the correction factor. The correction factor should be close to one for large samples and noticeably smaller than one for small samples. If the correction factor is not in the expected range, the degrees of freedom may be incorrect.

Fifth, verify the final effect size. The effect size should be consistent with the raw difference between the group means and the pooled standard deviation. If the effect size is not in the expected range, the calculation should be checked.

Common Failure Patterns in Effect Size Reporting

The most common failure pattern in effect size reporting is the use of Cohen's d in small samples without any acknowledgment of the bias. This failure is common because Cohen's d is the default metric in many statistical software packages, and researchers may not be aware of the bias or the correction.

A second common failure pattern is the inconsistent use of effect size metrics across studies. A researcher may report Cohen's d in one study and Hedges' g in another study, without explaining the difference. This inconsistency makes it difficult to compare the results across studies.

A third common failure pattern is the failure to report a confidence interval for the effect size. The confidence interval provides essential information about the precision of the effect size estimate, and its absence makes it difficult to interpret the effect size.

A fourth common failure pattern is the failure to report the sample sizes and the degrees of freedom. Without this information, the reader cannot verify the effect size calculation or determine whether the Hedges' correction was applied.

A fifth common failure pattern is the use of the effect size metric that is not appropriate for the study design. For example, a researcher may use Cohen's d for a study with unequal group sizes, even though the Hedges' correction is more appropriate for this design.

Limitations of Standardized Effect Sizes

Standardized effect sizes have several limitations that researchers should understand. The first limitation is that the standardized effect size is a relative measure, and it does not provide information about the absolute magnitude of the effect. A large standardized effect size can be associated with a small absolute difference if the variability is small, and a small standardized effect size can be associated with a large absolute difference if the variability is large.

The second limitation is that the standardized effect size depends on the variability of the sample. If the sample is more variable than the population, the standardized effect size will be smaller than the true effect. If the sample is less variable than the population, the standardized effect size will be larger than the true effect.

The third limitation is that the standardized effect size is not directly comparable across studies with different designs. A standardized effect size from a study with a paired design is not directly comparable to a standardized effect size from a study with an independent groups design.

The fourth limitation is that the standardized effect size does not provide information about the clinical or biological significance of the effect. A statistically significant effect size may not be biologically meaningful, and a biologically meaningful effect may not be statistically significant.

The fifth limitation is that the standardized effect size is sensitive to outliers. A single outlier can have a substantial effect on the sample mean and the sample standard deviation, which can affect the standardized effect size.

Safety and Regulatory Context for Effect Size Reporting

The choice between Cohen's d and Hedges' g has implications for the reporting of research results, and it is relevant to the broader context of research integrity and reproducibility. The Committee on Publication Ethics core practices emphasize the importance of accurate and transparent reporting of research results. The choice of effect size metric is part of this transparency, and the researcher should document the choice and the rationale.

The EQUATOR Network provides reporting guidelines for a wide range of study designs. These guidelines often include recommendations for the reporting of effect sizes, and the researcher should consult the relevant guideline for the study design.

The National Library of Medicine provides access to authoritative biomedical books and research-method references. These references can provide additional guidance on the choice and the reporting of effect sizes.

The National Institutes of Health provides guidance on the conduct of research and the reporting of results. The NIH Data Management and Sharing Policy emphasizes the importance of sharing data and the documentation of the analysis methods.

The ORCID provides a system for the identification of researchers and the maintenance of their research records. The researcher should maintain an accurate record of their research, including the effect size metrics used in their studies.

Professional Escalation Criteria

The researcher should seek professional advice when the effect size calculation is not straightforward. The following situations warrant consultation with a biostatistician or a research methodologist.

The first situation is when the data do not meet the assumptions of the effect size calculation. If the group variances are substantially different, or if the data are not normally distributed, the researcher should consult a biostatistician.

The second situation is when the study design is complex. The effect size calculation for a study with multiple groups, repeated measures, or a nested design is more complex than the calculation for a simple two-group design, and the researcher should consult a biostatistician.

The third situation is when the researcher is planning a meta-analysis. The meta-analysis requires the combination of effect sizes from multiple studies, and the researcher should consult a biostatistician to ensure that the effect sizes are comparable.

The fourth situation is when the researcher is uncertain about the appropriate effect size metric for the study design. The researcher should consult a biostatistician to ensure that the choice of metric is appropriate.

The fifth situation is when the researcher is reporting the effect size for a regulatory submission. The regulatory submission may have specific requirements for the reporting of effect sizes, and the researcher should consult the regulatory guidance.

A Decision Framework for Effect Size Selection Across Study Lifecycle Stages

The choice between Cohen's d and Hedges' g is not a single decision made once at the analysis stage. It is a sequence of decisions that should be revisited at each phase of a research project, from planning through peer review. A structured decision framework helps researchers avoid the common failure of defaulting to whatever metric their software outputs, and it ensures that the choice is defensible when reviewers or meta-analysts examine the work.

Stage One: Protocol Development and Preregistration

The first decision point occurs before any data are collected. At this stage, the researcher should estimate the expected sample size per group based on the study design, the availability of biological material, and the ethical constraints on subject numbers. This estimate is the primary input for the initial metric selection.

For a planned study with fewer than 20 subjects per group, the protocol should specify Hedges' g as the primary effect size metric. For a planned study with more than 50 subjects per group, either metric is acceptable, and the choice can be based on the reporting conventions in the field. For the intermediate range of 20 to 50 subjects per group, the researcher should compute the expected correction factor and decide whether the magnitude of the correction is practically meaningful for the research question.

The protocol should also specify whether the study will report the effect size with a confidence interval. The EQUATOR Network provides reporting guidelines for a wide range of study designs, and many of these guidelines include specific recommendations for effect size reporting. The researcher should consult the relevant guideline during the planning stage and incorporate its requirements into the protocol.

The protocol should be registered or archived before data collection begins. The National Institutes of Health requires that certain types of research be registered, and the Data Management and Sharing Policy emphasizes the importance of documenting analysis methods. A preregistration that states the planned effect size metric provides a record that can be compared against the final analysis to ensure that the metric was not changed after seeing the results.

Step Two: Data Collection and Interim Monitoring

During data collection, the researcher should monitor the actual sample sizes as they accumulate. The planned sample size may differ from the actual sample size due to animal loss, equipment failure, or other practical issues. If the actual sample size falls below the planned sample size, the researcher should reassess the choice of effect size metric.

For example, a study planned with 25 subjects per group may end with 18 subjects per group due to unexpected mortality. The planned sample size suggested that either metric was acceptable, but the actual sample size falls into the range where the Hedges' correction is meaningful. In this situation, the researcher should switch to Hedges' g and document the change in the methods section.

The researcher should also monitor the variability of the data as it accumulates. If the group variances are substantially different, the pooled standard deviation may not be an appropriate measure of variability. In this case, the researcher should consider whether a different effect size metric is more appropriate, and the decision should be documented.

Step Three: Analysis and Computation

At the analysis stage, the researcher computes the effect size using the metric selected in the protocol. The computation should be performed using the actual sample sizes, not the planned sample sizes. The researcher should verify that the software output matches the expected metric and that the correction factor is applied correctly.

The researcher should compute both Cohen's d and Hedges' g as a sensitivity check, even if only one metric is reported. This check allows the researcher to quantify the magnitude of the correction and to confirm that the choice of metric does not change the substantive interpretation of the results. If the difference between the two metrics is large enough to change the interpretation, the researcher should report both values and explain the difference.

The analysis code should be saved and documented. The code should specify the formula used to compute the effect size, the degrees of freedom, and the correction factor. This documentation allows the analysis to be reproduced by other researchers and allows the meta-analyst to combine the results with other studies.

Step Four: Reporting and Interpretation

The results section should report the effect size with a confidence interval, and it should state the metric that was used. The results section should also report the sample sizes, the means, and the standard deviations for each group, so that the reader can verify the effect size calculation.

The interpretation of the effect size should be based on the reported metric. A Hedges' g of 0.5 from a study with 10 subjects per group is not directly comparable to a Cohen's d of 0.5 from a study with 100 subjects per group. The researcher should interpret the effect size in the context of the sample size and the precision of the estimate.

The researcher should also consider whether the effect size is biologically meaningful. A statistically significant effect size may not be biologically meaningful, and a biologically meaningful effect may not be statistically significant. The researcher should discuss the biological significance of the effect size in the discussion section.

Step Five: Peer Review and Revision

The choice of effect size metric may be questioned during peer review. The reviewer may ask why the researcher chose Hedges' g instead of Cohen's d, or why the researcher did not report a confidence interval. The researcher should be prepared to answer these questions with the documentation from the protocol and the analysis.

The researcher should also be prepared to convert between Cohen's d and Hedges' g if the reviewer requests a different metric. The conversion is straightforward, and the researcher should provide the converted value with the correction factor and the degrees of freedom.

The Committee on Publication Ethics core practices emphasize the importance of accurate and transparent reporting of research results. The choice of effect size metric is part of this transparency, and the researcher should be prepared to justify the choice in the peer review process.

A Record System for Effect Size Decisions

The decision framework should be supported by a record system that documents the choice of effect size metric at each stage of the study lifecycle. The record system should include the following components:

The first component is the study protocol, which should state the planned sample size and the planned effect size metric. The protocol should be dated and versioned, and any changes to the protocol should be documented.

The second component is the data collection log, which should record the actual sample sizes as they accumulate. The log should also record any deviations from the planned sample size and the reason for the deviation.

The third component is the analysis code, which should include the formula used to compute the effect size. The code should be commented to explain the choice of metric and the correction factor.

The fourth component is the results section, which should report the effect size with a confidence interval and state the metric that was used. The results section should also report the sample sizes, the means, and the standard deviations for each group.

The fifth component is the revision log, which should document any changes to the effect size metric during the peer review process. The revision log should include the reviewer's comments and the researcher's response.

Troubleshooting Common Decision Errors

The decision framework can fail in several ways. The first failure pattern is the use of the planned sample size instead of the actual sample size. The researcher may plan for 30 subjects per group and select Cohen's d, but the actual sample size may be 15 per group. The effect size should be computed using the actual sample size, and the metric should be selected based on the actual sample size.

The second failure pattern is the failure to document the decision. The researcher may select Hedges' g but fail to state the choice in the methods section. This failure makes it difficult for the reader to verify the analysis and for the meta-analyst to combine the results with other studies.

The third failure pattern is the failure to report the confidence interval. The confidence interval provides essential information about the precision of the effect size estimate, and its absence makes it difficult to interpret the effect size.

The fourth failure pattern is the use of the effect size metric that is not appropriate for the study design. For example, a researcher may use Cohen's d for a study with unequal group sizes, even though the Hedges' correction is more appropriate for this design.

The fifth failure pattern is the failure to document the change in the effect size metric. The researcher may switch from Cohen's d to Hedges' g after seeing the results, but the change is not documented. This failure is a form of selective reporting, and it is a violation of the Committee on Publication Ethics core practices.

Escalation Criteria for the Decision Framework

The researcher should seek professional advice when the decision framework does not provide a clear answer. The following situations warrant consultation with a biostatistician or a research methodologist.

The first situation is when the actual sample size is much smaller than the planned sample size. The researcher may have planned for 30 subjects per group but ended with 8 subjects per group. The effect size calculation is more complex in this situation, and the researcher should consult a biostatistician.

The second situation is when the data do not meet the assumptions of the effect size calculation. If the group variances are substantially different, or if the data are not normally distributed, the researcher should consult a biostatistician.

The third situation is when the study design is complex. The effect size calculation for a study with multiple groups, repeated measures, or a nested design is more complex than the calculation for a simple two-group design, and the researcher should consult a biostatistician.

The fourth situation is when the researcher is planning a meta-analysis. The meta-analysis requires the combination of effect sizes from multiple studies, and the researcher should consult a biostatistician to ensure that the effect sizes are comparable.

The fifth situation is when the researcher is reporting the effect size for a regulatory submission. The regulatory submission may have specific requirements for the reporting of effect sizes, and the researcher should consult the regulatory guidance.

Integration with Research Records and Identity

The decision framework should be integrated with the researcher's broader record-keeping practices. The ORCID provides a system for the identification of researchers and the maintenance of their research records. The researcher should maintain an accurate record of their research, including the effect size metrics used in their studies.

The National Library of Medicine provides access to authoritative biomedical books and research-method references. These references can provide additional guidance on the choice and the reporting of effect sizes, and the researcher should consult these references when the decision framework is not clear.

The National Institutes of Health provides guidance on the conduct of research and the reporting of results. The NIH Data Management and Sharing Policy emphasizes the importance of sharing data and the documentation of the analysis methods. The researcher should ensure that the effect size decisions are documented in the data management plan and in the shared data.

Frequently Asked Questions

What is the main difference between Cohen's d and Hedges' g?

Cohen's d uses the pooled standard deviation of the two groups without any correction for sample size. Hedges' g applies a correction factor that adjusts for the bias in the standard deviation estimate when the sample size is small. The correction factor is always less than one, so Hedges' g is always slightly smaller than Cohen's d for the same data.

When should I use Hedges' g instead of Cohen's d?

Use Hedges' g when the total sample size is less than 50, which is common in biological research. The bias in Cohen's d is most pronounced at small sample sizes, and the Hedges' correction provides a more accurate estimate. For larger samples, the two metrics produce nearly identical values, and the choice has no practical consequence.

Does the choice between Cohen's d and Hedges' g affect my statistical significance?

The choice between Cohen's d and Hedges' g does not affect the p-value or the statistical significance of the test. The effect size is a descriptive statistic that describes the magnitude of the difference between groups, while the p-value is a measure of the evidence against the null hypothesis. The two metrics are computed from the same data, and the choice between them does not change the p-value.

Can I convert Cohen's d to Hedges' g?

Yes, Hedges' g can be computed from Cohen's d by multiplying Cohen's d by the correction factor. The correction factor depends on the total degrees of freedom, which is the sum of the two group sample sizes minus two. The correction factor is close to one for large samples and noticeably smaller than one for small samples.

Why does Cohen's d overestimate the effect size in small samples?

Cohen's d overestimates the effect size in small samples because the sample standard deviation tends to underestimate the true standard deviation when the sample size is small. The underestimated standard deviation makes the denominator in the effect size formula smaller, which makes the effect size larger. The Hedges' correction adjusts for this bias.

What is the pooled standard deviation in the effect size formula?

The pooled standard deviation is the square root of the weighted average of the individual group variances. The weights are the degrees of freedom from each group. The pooled standard deviation is used as the denominator in the effect size formula, and it represents the common variability of the two groups.

Should I report the effect size with a confidence interval?

Yes, the effect size should be reported with a confidence interval. The confidence interval provides information about the precision of the effect size estimate, and it allows the reader to assess whether the effect size is consistent with a meaningful biological effect. A wide confidence interval indicates that the effect size is not precisely estimated.

How do I document the choice of effect size metric in my study?

Document the choice of effect size metric in the study plan, the analysis methods, and the results. The documentation should state the sample size, the degrees of freedom, and the formula used to compute the effect size. This documentation allows the researcher to reproduce the analysis and allows the reader to understand the analysis.

Using the Evidence

SourceBest use in this topicImportant limitation
Research Methods Resourcesofficial guidanceCheck the linked page for current local requirements
EQUATOR Networkofficial guidanceCheck the linked page for current local requirements
Core Practicesofficial guidanceCheck the linked page for current local requirements

Related Bioinformatics Guides

Related Clinical & Scientific Guides

References and Further Reading

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.