Guides / forest plot
Forest plots: how to read and interpret them
A forest plot puts every study in a meta-analysis on one chart alongside the pooled result, and it carries more information than most readers extract from it. This guide explains each element, shows how to judge heterogeneity by eye, and sets out the misreadings that most often appear in discussion sections.
A forest plot is a graph used in meta-analysis that displays the effect estimate and confidence interval for every included study on a common scale, together with the pooled result. Each study appears as a square positioned at its effect estimate, sized by the weight it carries, with a horizontal line showing its confidence interval. The pooled estimate is shown as a diamond at the bottom.
Definition
What a forest plot shows
A forest plot displays every study in a meta-analysis on a single scale, with the pooled result beneath them. It is the standard way of presenting a meta-analysis because it shows the individual evidence and the summary at the same time, rather than replacing one with the other.
The name comes from the appearance of the plot: a row of vertical lines and squares resembling trees, against the vertical line of no effect. Its value is that it makes visible what a pooled number alone conceals — whether the studies agree, whether one dominates, and how precise each contribution is.
The abstract gives you a pooled number. The forest plot tells you whether that number is a fair summary or whether it is being driven by one large study, dragged by an outlier, or averaging over studies that flatly disagree.
Elements
The anatomy of a forest plot
| Element | What it represents |
|---|---|
| Study label | Usually first author and year, sometimes with the country or design |
| Square | The study’s point estimate; its area is proportional to weight |
| Horizontal line | The 95% confidence interval for that study |
| Arrowhead on a line | The interval extends beyond the plotted axis |
| Vertical line | The null value — 0 for differences, 1 for ratios |
| Diamond | The pooled estimate; its width is the pooled confidence interval |
| Numeric column | The same estimates and intervals as text |
| Weight column | Percentage contribution of each study |
For ratio measures — odds ratios, risk ratios, hazard ratios — the axis is logarithmic and the null line sits at 1, not 0. On a log axis, an effect of 0.5 and an effect of 2.0 are the same distance from the null in opposite directions, which is exactly why the log scale is used. Reading a ratio plot as though the axis were linear will mislead you about relative magnitude.
Interpretation
Reading it row by row
It is worth pausing on why the plot is laid out this way at all. A table of estimates and intervals contains identical information, and for a handful of studies is arguably easier to read precisely. What the plot adds is the ability to take in agreement and disagreement at a glance across twenty or thirty rows, where a table would require the reader to hold each interval in mind while comparing it against the others. That is the whole justification for the format, and it explains why the elements are chosen as they are — every one of them encodes something you would otherwise have to compute.
For each study, three things matter, and they are worth taking in order.
Applying this to the plot above: Ahmed 2021 sits at 0.30 with a short interval clear of the null — a precise, significant positive result. Dubois 2023 sits at 0.28, almost the same estimate, but its interval runs from −0.36 to 0.92 and crosses the null. Those two studies agree closely on the size of the effect; they differ entirely in how confidently they establish it.
Dubois 2023 is frequently described as a study that “found no effect”. It found an effect of 0.28 — almost identical to the largest study in the plot — but lacked the precision to exclude zero. Treating it as evidence against the intervention misreads the plot.
Weighting
Why the squares are different sizes
Square area is proportional to the weight the study carries in the pooled estimate, and weight is determined by precision: 1 / variance. Since variance falls as sample size rises, larger studies get bigger squares.
| Study | Weight | What that means |
|---|---|---|
| Ahmed 2021 | 42.1% | Contributes more than the bottom three combined |
| Bianchi 2022 | 24.3% | Substantial |
| Costa 2020 | 14.8% | Moderate |
| Dubois 2023 | 10.5% | Limited influence |
| Evans 2019 | 8.3% | Minimal influence |
The visual purpose of the squares is to stop your eye from treating all rows equally. Without them, five studies look like five equal votes; with them, it is immediately clear that one study accounts for over 40% of the result.
Under a fixed effect model, weights are driven purely by precision, so large studies dominate heavily. Under random effects, the between-study variance is added to every study's variance before weighting, which compresses the differences and gives small studies relatively more influence. The same five studies can produce visibly different squares depending only on the model chosen.
Our effect size calculator turns means, SDs and group sizes into Cohen's d and Hedges' g with confidence intervals — the inputs a forest plot needs.
Use the calculatorThe pooled result
The diamond, and the line of no effect
The diamond represents the pooled estimate. Its centre is the point estimate and its left and right points mark the confidence interval, so a narrow diamond means a precisely estimated pooled effect.
| Diamond position | Interpretation |
|---|---|
| Entirely to the right of the null | Significant effect favouring the intervention |
| Entirely to the left | Significant effect favouring the control |
| Touching or crossing the null | Pooled effect not statistically significant |
In the plot above the diamond sits at 0.39 with an interval of [0.24, 0.54], clear of the null line — a significant pooled effect of moderate size.
The diamond shape is not decorative. Because its widest point is at the centre and it tapers to the interval bounds, it draws the eye to the point estimate while still showing the range — and it is visually distinct from the study rows above, so nobody mistakes the summary for another study. Some plots additionally draw a vertical dashed line through the diamond's centre, extended up through the study rows, which makes it easy to see at a glance which individual studies sat above and below the pooled result.
Ordering matters more than it appears. Studies are most often listed alphabetically or by year, neither of which helps the reader. Sorting by effect size makes any gradient immediately visible; sorting by weight puts the influential studies at the top where they belong. Cochrane reviews conventionally use year order to show how the evidence accumulated, which is defensible — but whichever you choose, say so in the caption, because a reader who assumes one ordering and gets another will misread the pattern.
Which side is which
Always read the axis labels beneath the plot before interpreting direction. “Favours intervention” may be on the left or the right depending on how the outcome was coded, and for harmful outcomes such as mortality a lower value is the better result. Assuming right means good is a genuine and common source of error.
A second, wider line through the diamond shows the prediction interval — the range in which the effect of a future study would be expected to fall. Where heterogeneity is substantial, the diamond can be clear of the null while the prediction interval crosses it, meaning the average effect is positive but a new study might find nothing. That distinction is usually the most practically important thing on the chart.
Heterogeneity
Spotting heterogeneity by eye
Before reading the statistics, look at how the intervals line up.
Low heterogeneity
- Intervals overlap substantially
- Estimates cluster in a narrow band
- All studies point the same direction
- I² typically below 40%
High heterogeneity
- Some intervals barely overlap
- Estimates scattered widely
- Studies point in opposite directions
- I² typically above 60%
A useful habit is to look at the plot twice: once ignoring the diamond entirely, to form your own impression of what the studies show, and once with it. If your impression from the individual rows differs noticeably from the pooled result, that gap is worth understanding before citing the number. Usually it means one heavily weighted study is doing most of the work, or that a couple of imprecise studies pointing the other way are contributing almost nothing to the average despite occupying as much vertical space as the rest.
Statistics accompanying the plot usually appear beneath it as Q, I² and τ². Read them alongside the visual impression rather than instead of it: I² is a proportion and can be high simply because the studies were precise, so a plot of tightly estimated studies with small genuine differences may report a large I² that overstates the practical disagreement.
Judging by eye has a known weakness worth naming: the visual impression is driven by how wide the plotted axis happens to be. Software chooses that range to fit the most extreme interval, so a single very imprecise study stretches the axis and makes every other interval look narrow and tightly clustered. Two plots of the same data can therefore give quite different impressions of consistency. Check the axis values before trusting the visual impression.
A pooled estimate averaging a strong positive and a strong negative finding describes neither. If the forest plot shows genuine disagreement in direction, the productive response is to investigate why — through subgroup analysis or meta-regression — rather than to report the average as though it settled the question.
Variants
Subgroup and sensitivity plots
Where an interval extends beyond the plotted axis, software draws an arrowhead rather than rescaling the whole chart around one imprecise study. That is the right visual choice, but it hides the true extent of that study's uncertainty. Always read the numeric column alongside the graphic, because a row with an arrow is telling you the study established very little.
| Variant | What it shows | Use |
|---|---|---|
| Subgroup forest plot | Studies grouped, each with its own diamond | Testing whether the effect differs by a study characteristic |
| Cumulative forest plot | Studies added one at a time, usually by year | Showing how the evidence accumulated over time |
| Leave-one-out plot | Pooled estimate with each study removed in turn | Testing whether one study drives the result |
| Multi-outcome plot | Several outcomes stacked | Summarising a review with several endpoints |
A subgroup plot shows a diamond per subgroup plus an overall diamond, and is usually accompanied by a test for subgroup differences. Note that subgroup analyses are observational even within a randomised trial — the studies were not randomly assigned to subgroups — so a difference between them is a hypothesis, not a causal finding, and it must be pre-specified to carry much weight.
Leave-one-out plots deserve wider use. If removing a single study moves the pooled estimate substantially or changes its significance, the conclusion rests on that study and the review should say so plainly.
Send your extraction table. A named statistician runs the meta-analysis and returns forest, funnel and sensitivity plots with the heterogeneity properly assessed.
See statistical consultancyPitfalls
Six misreadings to avoid
1. Treating every row as an equal vote
Counting how many studies were significant ignores the weights entirely. Look at the square sizes.
2. Reading a wide interval as a contradictory finding
An imprecise study that agrees on the estimate but crosses the null is not evidence against the effect.
3. Assuming right means favourable
Read the axis labels. For outcomes such as mortality or relapse, a result on the left is the better one.
4. Ignoring the scale on ratio plots
Odds and risk ratios are plotted on a log axis with the null at 1. Distances are not linear.
5. Reporting the diamond without the heterogeneity
A pooled estimate from studies that disagree needs that disagreement reported alongside it.
6. Overlooking a dominant study
If one study carries 60% of the weight, the meta-analysis is largely reporting that study. Run a leave-one-out analysis and say so.
Reporting
Producing and reporting a forest plot
Blank rows and section headings earn their space in longer plots. Grouping studies by design, setting or risk of bias with a labelled heading between the groups costs nothing and reveals structure that would otherwise be buried in an undifferentiated list of thirty rows. Most meta-analysis packages support this directly.
One practical caution on producing your own: the plot must be generated from the same analysis that produced your reported numbers, not assembled separately. Hand-drawn or spreadsheet-built plots drift out of step with the analysis as soon as a study is added or a model changed, and a figure that disagrees with the text is the kind of inconsistency reviewers notice immediately.
Standard software includes the metafor and meta packages in R, the metan suite in Stata, and RevMan for Cochrane reviews. All produce publication-standard plots directly; drawing one by hand in a graphics program is both laborious and prone to introducing errors of scale.
“Forest plot of five studies (N = 1,758) comparing retrieval practice with conventional revision on end-of-module performance. Random effects model. Squares are proportional to study weight; the diamond shows the pooled estimate. Pooled SMD = 0.39, 95% CI [0.24, 0.54]; Q(4) = 3.12, p = .538, I² = 0%.”
Answers
Frequently asked questions
What is a forest plot used for?
To display the results of every study in a meta-analysis on a single scale alongside the pooled estimate. It lets a reader see each study's effect and precision, how much weight each carries, and whether the studies agree — information a pooled number alone cannot convey.
What does the diamond on a forest plot mean?
It represents the pooled estimate from the meta-analysis. Its centre is the point estimate and its left and right points mark the confidence interval. If the diamond touches or crosses the line of no effect, the pooled result is not statistically significant.
Why are the squares different sizes?
Square area is proportional to the weight each study carries, which is determined by its precision — the inverse of its variance. Larger, more precise studies get bigger squares and influence the pooled estimate more. The sizing exists to stop readers treating every study as an equal vote.
What does it mean when a confidence interval crosses the line of no effect?
That study did not reach statistical significance on its own. It does not mean the study found no effect — the point estimate may be substantial — only that the interval is wide enough to include the possibility of no effect, usually because the study was small.
How do I tell if there is heterogeneity from a forest plot?
Look at whether the confidence intervals overlap. Substantial overlap with estimates clustered together suggests low heterogeneity; scattered estimates with barely overlapping intervals, especially pointing in opposite directions, suggest high heterogeneity. Confirm against the I squared and tau squared reported beneath the plot.
Why is the axis logarithmic on some forest plots?
Because ratio measures such as odds ratios and risk ratios are not symmetric on a linear scale — halving and doubling should appear as equal distances from the null. A log axis achieves that, and the null line sits at 1 rather than 0.
What software makes forest plots?
The metafor and meta packages in R, the metan suite in Stata, and RevMan for Cochrane reviews, all of which produce publication-standard output directly. Drawing one manually in a graphics program risks introducing errors of scale and is not recommended.
What is the difference between a confidence interval and a prediction interval on a forest plot?
The confidence interval, shown by the diamond's width, describes how precisely the average effect has been estimated. The prediction interval, sometimes drawn as a wider line through the diamond, describes where the true effect in a new setting would be expected to fall. Where heterogeneity is substantial the prediction interval is much wider and is usually the more useful figure.
Send the data. Get a fixed quote.
Attach your dataset or just describe the project. A named statistician replies with a price and a deadline, usually within one working day.