# How to Read a Forest Plot in a Meta-Analysis

A forest plot is the signature graphic of a meta-analysis. Each row shows one study's effect estimate as a square, with a horizontal line for its confidence interval, and a diamond at the bottom shows the pooled result. The plot packs the entire quantitative argument of a systematic review into a single figure, which is why it appears in Cochrane reviews, clinical guidelines and journal articles across medicine.

If you read primary literature, you will meet forest plots constantly, and they are easy to misread. A square that looks impressive may carry almost no weight. A diamond that clears the line of no effect may sit on top of studies that disagree with each other. This guide explains what every element means, how the underlying statistics produce the picture, and where careful readers slow down.

## Quick Answer

- A forest plot displays, for each study, a point estimate (the square) and its confidence interval (the horizontal line), plus a pooled estimate (the diamond) [1].
- The area of each square reflects that study's weight in the meta-analysis, so bigger squares pull the diamond toward them [1].
- The vertical line of no effect sits at 1 for ratio measures (risk ratio, odds ratio) and at 0 for difference measures (mean difference) [2].
- If a study's confidence interval or the pooled diamond crosses that line, the result is not statistically significant at the 5% level [2].
- The pooled estimate under the inverse-variance method is a weighted average: $\hat{Y} = \frac{\sum_i Y_i W_i}{\sum_i W_i}$, where $Y_i$ is study $i$'s effect estimate and $W_i = 1 / SE_i^2$ is its weight [1].
- Heterogeneity statistics (Cochran's Q, $I^2$, $\tau^2$) tell you whether the studies are similar enough for one pooled number to mean anything [1].

## What Is a Forest Plot?

A forest plot is a graph that lines up the results of several studies of the same question. Each study gets a row. The square sits at the study's point estimate of the intervention effect, and the horizontal line either side of it shows the confidence interval [1]. The area of the square indicates the weight assigned to that study in the meta-analysis [1]. Because weight usually tracks precision, the block size draws your eye to the studies with larger weight, which dominate the summary result shown as a diamond at the bottom [1].

The diamond represents the pooled effect. Its center is the pooled estimate, and its width shows the confidence interval for the overall effect [2]. A vertical line marks the line of no effect: at 1 for ratio measures such as the risk ratio or odds ratio, and at 0 for difference measures such as a mean difference [2]. When the 95% confidence interval of a single study or of the pooled estimate crosses the line of no effect, that difference is not statistically significant at the 5% level [2].

The logic is simple: one row per study, one summary at the bottom, and a reference line to judge direction.

## How the Numbers Behind the Squares Are Built

For ratio measures, the analysis is done on the natural log scale. Cochrane specifies that data must be entered as log odds ratios (or log risk ratios) with the standard error of the log ratio [1]. Working in logs makes the sampling distribution closer to normal and makes symmetric confidence intervals possible.

Take a 2x2 table with cells $a$, $b$, $c$, $d$. The log odds ratio and its approximate variance (the Woolf method) are:

$$\ln(\text{OR}) = \ln\left(\frac{ad}{bc}\right), \qquad \text{Var}(\ln \text{OR}) \approx \frac{1}{a} + \frac{1}{b} + \frac{1}{c} + \frac{1}{d}$$

The standard error is the square root of that variance. A 95% confidence interval on the log scale is the estimate plus or minus 1.96 standard errors, and exponentiating the endpoints returns the interval to the odds ratio scale.

The inverse-variance method then takes a weighted average [1]:

$$\hat{Y} = \frac{\sum_i Y_i W_i}{\sum_i W_i}, \qquad W_i = \frac{1}{SE_i^2}$$

Here $Y_i$ is study $i$'s effect estimate and $W_i$ is its weight. Because weights are $1/SE^2$, studies with smaller standard errors, usually larger studies with more events, get more influence [1]. The standard error of the fixed-effect pooled estimate is $1 / \sqrt{\sum W_i}$, and its 95% confidence interval is the pooled estimate plus or minus 1.96 times that standard error.

Ratio axes are drawn on a log scale so that effects of equal size in opposite directions sit at equal distances from 1. The natural log of 2 is 0.693 and the natural log of 0.5 is -0.693, so an odds ratio of 2 and an odds ratio of 0.5 are mirror images around the line of no effect. On a log axis, a confidence interval for a ratio looks symmetric around the point estimate even though, on the raw scale, it is not. An odds ratio of 0.621 with a confidence interval of 0.436 to 0.884 is a good example: the raw-scale distances differ, but the log-scale distances match.

## Fixed Effects, Random Effects and the Diamond

A fixed-effect meta-analysis assumes all studies estimate the same underlying intervention effect [1]. A random-effects meta-analysis assumes the underlying effects follow a distribution across studies, and the between-study variance is called $\tau^2$ [1]. Borenstein and colleagues stress that the two models are not interchangeable: they rest on fundamentally different assumptions about the data [6].

Under any interpretation, a fixed-effect meta-analysis ignores heterogeneity [1]. In a heterogeneous set of studies, a random-effects meta-analysis gives relatively more weight to smaller studies than a fixed-effect analysis would [1]. The random-effects summary is an average intervention effect across studies, not an effect assumed to be identical in every study [1].

The simplest random-effects version of the inverse-variance method is the DerSimonian and Laird method from 1986 [1][5]. Its between-study variance is:

$$\tau^2 = \max\left(0, \frac{Q - df}{C}\right), \qquad C = \sum_i W_i - \frac{\sum_i W_i^2}{\sum_i W_i}$$

where $Q$ is Cochran's heterogeneity statistic and $df$ is the degrees of freedom. Random-effects weights become $1 / (SE_i^2 + \tau^2)$. In RevMan the default estimator of the between-study variance is REML, with the DerSimonian and Laird moment method still available [1]. Cochrane notes that DerSimonian and Laird is the simplest but not the best estimator; REML or other estimators can give different $\tau^2$ values and confidence intervals, so check the current documentation for whichever software you use.

Heterogeneity itself is judged in three ways on a forest plot: overlap of point estimates and confidence intervals, the Chi-squared P value, and $I^2$ [2]. Cochran's Q is:

$$Q = \sum_i W_i (Y_i - \hat{Y})^2$$

compared with a chi-squared distribution on $k - 1$ degrees of freedom, where $k$ is the number of studies. Poor overlap of the studies' confidence intervals generally indicates statistical heterogeneity [1]. The Chi-squared test has low power when studies are few or small, so a non-significant result must not be taken as evidence of no heterogeneity; a P value of 0.10 is sometimes used instead of 0.05 [1]. With many studies, the test can have high power to detect a small, clinically unimportant amount of heterogeneity [1].

$I^2$ converts Q into a percentage:

$$I^2 = \frac{Q - df}{Q} \times 100\%$$

It describes the percentage of variability in effect estimates that is due to heterogeneity and not to sampling error [1]. Higgins and Thompson, who proposed $H$, $R$ and $I^2$, note that $I^2$ does not depend on the number of studies or the effect metric, and they recommend presenting $H$ or $I^2$ in preference to the heterogeneity test [3]. The Cochrane rough guide is that 0% to 40% might not be important, 30% to 60% may be moderate, 50% to 90% may be substantial, and 75% to 100% is considerable heterogeneity [1]. Cochrane adds that the importance of an $I^2$ value depends on the magnitude and direction of effects and the strength of evidence for heterogeneity, that uncertainty in $I^2$ is substantial when there are few studies, and that simple thresholds should not be used to diagnose heterogeneity [1].

$\tau$ is the estimated standard deviation of underlying effects across studies. Cochrane recommends prediction intervals as a more interpretable way to express it [1]. The simple 95% prediction interval is:

$$M \pm t_{k-1} \times \sqrt{\tau^2 + SE(M)^2}$$

where $M$ is the random-effects mean, $t_{k-1}$ is the 97.5th percentile of the t distribution with $k - 1$ degrees of freedom, and $SE(M)$ is the standard error of $M$ [1]. Prediction intervals rely strongly on normally distributed study effects and can be very problematic when the number of studies is small [1].

## Worked Example

Five made-up trials compare a drug with control, and the outcome is a bad event. The table shows the raw counts, the per-study odds ratio with its Woolf standard error on the log scale, and the 95% confidence interval.

| Study | Events/Total (treatment) | Events/Total (control) | OR | ln OR | SE | 95% CI |
|---|---|---|---|---|---|---|
| A | 8/100 | 25/100 | 0.261 | -1.3437 | 0.4350 | 0.111 to 0.612 |
| B | 30/250 | 45/250 | 0.621 | -0.4761 | 0.2549 | 0.377 to 1.024 |
| C | 8/60 | 10/60 | 0.769 | -0.2624 | 0.5140 | 0.281 to 2.107 |
| D | 60/500 | 90/500 | 0.621 | -0.4761 | 0.1802 | 0.436 to 0.884 |
| E | 18/80 | 10/80 | 2.032 | 0.7091 | 0.4312 | 0.873 to 4.732 |

Fixed-effect inverse-variance weights, expressed as a percentage of the total, are A 8.7%, B 25.4%, C 6.2%, D 50.8% and E 8.9%. Study D would be drawn with the largest square because its standard error is smallest.

The fixed-effect pooled log odds ratio is -0.4333 with a standard error of 0.1284, giving a pooled odds ratio of 0.648 (95% CI 0.504 to 0.834), z = -3.37, p = 0.00074. The diamond sits left of 1 and does not touch it.

Heterogeneity is substantial. Q = 11.594 on 4 degrees of freedom, p = 0.021, and $I^2$ = (11.594 - 4) / 11.594 = 65.5%. The DerSimonian-Laird estimate is $\tau^2$ = 0.190, so $\tau$ = 0.436. Random-effects weights shift toward the smaller studies: A 16.6%, B 24.6%, C 13.8%, D 28.2%, E 16.7%. Study D loses weight, and the small studies gain it.

The random-effects pooled odds ratio is 0.676 (95% CI 0.413 to 1.104), so the wider random-effects diamond crosses the line of no effect. The 95% prediction interval, using the Cochrane formula with $t$ on 4 degrees of freedom equal to 2.776, runs from an odds ratio of 0.167 to 2.732. A new similar study could plausibly show benefit or harm.

Reading the plot: four studies favor treatment and Study E points the other way. The fixed-effect diamond suggests a clear benefit, but the $I^2$ of 65.5%, moderate to substantial on the Cochrane guide, warns that the studies disagree more than chance explains, and the random-effects result is uncertain.

## How to Read a Forest Plot Step by Step

Start with the axis and the line of no effect. Confirm whether the measure is a ratio or a difference, because that determines whether the reference line sits at 1 or 0 [2]. Check whether the axis is on a log scale; if it is, equal visual distances from 1 correspond to equal multiplicative effects in opposite directions.

Scan the squares and their sizes. The largest square carries the most weight, and the diamond tends to sit close to it [1]. A study with a dramatic point estimate but a tiny square contributes little.

Look at the horizontal lines. A line that crosses the line of no effect means that study alone is not statistically significant at the 5% level [2]. That is not proof of no effect. In the example, Study C has an odds ratio of 0.77 with a confidence interval of 0.28 to 2.11, which is compatible with a large benefit and with doubled odds of harm.

Find the diamond and check which model produced it. A fixed-effect diamond answers a narrow question about one common effect. A random-effects diamond answers a broader question about the average effect across a distribution of effects [1][6]. If the two diamonds differ in width or in position relative to the line of no effect, heterogeneity is doing real work.

Then read the heterogeneity statistics. Overlap of confidence intervals, the Chi-squared P value and $I^2$ together give a fuller picture than any one alone [2]. If $I^2$ is high, ask whether the studies are clinically similar enough for pooling to make sense, and look for the prediction interval if the authors report one [1].

## Common Mistakes

- Treating a confidence interval that crosses 1 as proof of no effect. It only means the result is not statistically significant at the 5% level [2]. The interval may still include clinically important benefits and harms.
- Reading the fixed-effect diamond as the final answer when heterogeneity is high. In the example, the fixed-effect odds ratio of 0.648 (0.504 to 0.834) looks decisive, but with $I^2$ = 65.5% the random-effects odds ratio of 0.676 (0.413 to 1.104) crosses 1. The fixed-effect result can hide disagreement among studies.
- Judging a study by the height of its point estimate instead of its weight. The square area encodes weight, and a small square means little influence on the diamond [1].
- Assuming a non-significant Chi-squared test rules out heterogeneity. The test has low power when studies are few or small [1].
- Using simple $I^2$ cutoffs as a diagnosis. Cochrane states that the importance of an $I^2$ value depends on the magnitude and direction of effects and the strength of evidence for heterogeneity, and that simple thresholds should not be used to diagnose heterogeneity [1].
- Ignoring the scale. On a log axis, an odds ratio of 0.5 and an odds ratio of 2 are equally far from 1, which is not true on a raw axis.

## Limitations

The Woolf variance for the log odds ratio fails when any cell in the 2x2 table is zero. Continuity corrections, or methods such as Mantel-Haenszel or Peto, are used in that situation [1]. Check which method a review used before comparing its numbers with yours.

The choice of random-effects estimator changes the result. DerSimonian and Laird is transparent and easy to compute by hand, but Cochrane notes it is not the best available estimator, and REML or other approaches can give different $\tau^2$ values and confidence intervals [1]. The worked example uses DerSimonian and Laird for that transparency, not because it is the default everywhere.

Prediction intervals depend on an assumption of normally distributed study effects and behave poorly when the number of studies is small [1]. The degrees of freedom used for the t multiplier have also varied across Handbook versions, so intervals computed with different conventions will not match exactly. Check the current documentation for the software and Handbook version you are using.

Interpreting a confidence interval that crosses 1 as "not significant" is a statement about the 5% level only. It says nothing about the absence of an effect, and it does not capture the full uncertainty in a heterogeneous body of evidence.

## Frequently Asked Questions

### What is a forest plot in simple terms?

It is a graph that shows one row per study, with a square for the point estimate and a line for the confidence interval, plus a diamond for the pooled result [1]. A vertical line marks no effect, at 1 for ratios and 0 for differences [2]. The size of each square shows how much that study contributed to the pooled estimate.

### How do I interpret a forest plot when the diamond crosses the line of no effect?

The pooled result is not statistically significant at the 5% level [2]. That does not mean the treatment does nothing; the confidence interval still spans a range of possible effects, some of which may matter clinically. Look at the width of the diamond and the heterogeneity statistics before drawing a conclusion.

### What does the diamond in a forest plot represent?

The diamond is the pooled effect. Its center is the pooled estimate and its width is the confidence interval for the overall effect [2]. A wide diamond means an imprecise summary, often because studies are small or because a random-effects model was used in the presence of heterogeneity.

### What is a good heterogeneity I2 value?

Cochrane's rough guide is that 0% to 40% might not be important, 30% to 60% may be moderate, 50% to 90% may be substantial, and 75% to 100% is considerable [1]. Cochrane also warns that the importance of an $I^2$ value depends on the magnitude and direction of effects and the strength of evidence for heterogeneity, and that simple thresholds should not be used to diagnose heterogeneity [1]. Treat the ranges as orientation, not as a pass or fail test.

### Why do some forest plots use a log scale?

For ratio measures, the analysis is done on the natural log scale, and Cochrane specifies entering data as log odds ratios or log risk ratios with the standard error of the log ratio [1]. A log axis places effects of equal size in opposite directions at equal distances from 1, so an odds ratio of 2 and an odds ratio of 0.5 mirror each other. It also makes confidence intervals for ratios look symmetric around the point estimate.

## References

1. [Cochrane Handbook, Chapter 10: Analysing data and undertaking meta-analyses (v6.5)](https://www.cochrane.org/authors/handbooks-and-manuals/handbook/current/chapter-10)
2. [Dettori JR, Norvell DC, Chapman JR. How to interpret a meta-analysis forest plot. Global Spine J 2021;11:614-616](https://doi.org/10.1177/21925682211003889)
3. [Higgins JPT, Thompson SG. Quantifying heterogeneity in a meta-analysis. Stat Med 2002;21:1539-1558](https://doi.org/10.1002/sim.1186)
4. [Higgins JPT et al. Measuring inconsistency in meta-analyses. BMJ 2003;327:557-560](https://doi.org/10.1136/bmj.327.7414.557)
5. [DerSimonian R, Laird N. Meta-analysis in clinical trials. Control Clin Trials 1986;7:177-188](https://doi.org/10.1016/0197-2456(86)90046-2)
6. [Borenstein M et al. Fixed-effect and random-effects models for meta-analysis. Res Synth Methods 2010;1:97-111](https://doi.org/10.1002/jrsm.12)
7. [Lewis S, Clarke M. Forest plots: trying to see the wood and the trees. BMJ 2001;322:1479-1480](https://doi.org/10.1136/bmj.322.7300.1479)

## Related Articles

- [How to Report a Meta-Analysis](/blog/research-skills/how-to-report-a-meta-analysis-a-checklist-for-transparent-and-complete-results)
- [Understanding Heterogeneity in Meta-Analysis](/blog/research-skills/understanding-heterogeneity-in-meta-analysis-how-to-interpret-i2-tau2-and-q-statistics)
- [Credible Intervals vs. Confidence Intervals](/knowledge/bioinformatics/credible-intervals-vs-confidence-intervals-what-biologists-need-to-know-for-correct-interpretation)
- [Full-Text Review in Systematic Reviews](/blog/research-skills/full-text-review-in-systematic-reviews-how-to-manage-the-screening-stage-and-avoid-common-pitfalls)
- [Absolute Risk Reduction vs Relative Risk Reduction: Meaning, Formulas and Examples](/blog/research-skills/absolute-risk-reduction-vs-relative-risk-reduction)
- [Hazard Ratio: What It Means, How to Interpret It, and How It Differs From Odds Ratio and Relative Risk](/blog/research-skills/hazard-ratio-explained)