Guides  /  forest plot

Forest plots: how to read and interpret them

A forest plot puts every study in a meta-analysis on one chart alongside the pooled result, and it carries more information than most readers extract from it. This guide explains each element, shows how to judge heterogeneity by eye, and sets out the misreadings that most often appear in discussion sections.

Hafiz Ahmad Tariq Written and reviewed by Hafiz Ahmad Tariq, Senior Biostatistician
Updated 17 August 202616 min read
What is a forest plot?

A forest plot is a graph used in meta-analysis that displays the effect estimate and confidence interval for every included study on a common scale, together with the pooled result. Each study appears as a square positioned at its effect estimate, sized by the weight it carries, with a horizontal line showing its confidence interval. The pooled estimate is shown as a diamond at the bottom.

Definition

What a forest plot shows

A forest plot displays every study in a meta-analysis on a single scale, with the pooled result beneath them. It is the standard way of presenting a meta-analysis because it shows the individual evidence and the summary at the same time, rather than replacing one with the other.

The name comes from the appearance of the plot: a row of vertical lines and squares resembling trees, against the vertical line of no effect. Its value is that it makes visible what a pooled number alone conceals — whether the studies agree, whether one dominates, and how precise each contribution is.

confidence intervals, a line of no effect and a diamond for the pooled estimate"> Study SMD [95% CI] Weight line of no effect Ahmed 2021Bianchi 2022Costa 2020 Dubois 2023Evans 2019 0.30 [0.10, 0.50]0.45 [0.18, 0.72] 0.62 [0.21, 1.03]0.28 [-0.36, 0.92] 0.55 [0.05, 1.05] 42.1%24.3%14.8% 10.5%8.3% Pooled (random) 0.39 [0.24, 0.54] 100% −0.500.5 ← favours control favours intervention →
A forest plot with five studies. Square size shows weight, the horizontal line shows the confidence interval, and the diamond at the bottom is the pooled estimate.
Read the plot before the abstract

The abstract gives you a pooled number. The forest plot tells you whether that number is a fair summary or whether it is being driven by one large study, dragged by an outlier, or averaging over studies that flatly disagree.

Elements

The anatomy of a forest plot

ElementWhat it represents
Study labelUsually first author and year, sometimes with the country or design
SquareThe study’s point estimate; its area is proportional to weight
Horizontal lineThe 95% confidence interval for that study
Arrowhead on a lineThe interval extends beyond the plotted axis
Vertical lineThe null value — 0 for differences, 1 for ratios
DiamondThe pooled estimate; its width is the pooled confidence interval
Numeric columnThe same estimates and intervals as text
Weight columnPercentage contribution of each study
Check which scale you are on

For ratio measures — odds ratios, risk ratios, hazard ratios — the axis is logarithmic and the null line sits at 1, not 0. On a log axis, an effect of 0.5 and an effect of 2.0 are the same distance from the null in opposite directions, which is exactly why the log scale is used. Reading a ratio plot as though the axis were linear will mislead you about relative magnitude.

Interpretation

Reading it row by row

crosses the line → NOT significant clear of the line → significant wide interval = imprecise study null value (0 for differences, 1 for ratios) Square size = the study’s WEIGHT. Line length = its PRECISION. They are not the same thing. A big square with a short line is a large, precise study — the rows that drive the pooled result.
Square size and line length carry different information. Weight and precision are related but not identical.

It is worth pausing on why the plot is laid out this way at all. A table of estimates and intervals contains identical information, and for a handful of studies is arguably easier to read precisely. What the plot adds is the ability to take in agreement and disagreement at a glance across twenty or thirty rows, where a table would require the reader to hold each interval in mind while comparing it against the others. That is the whole justification for the format, and it explains why the elements are chosen as they are — every one of them encodes something you would otherwise have to compute.

For each study, three things matter, and they are worth taking in order.

Where is the square? That is the study’s best estimate of the effect
How long is the line? A short line means a precise estimate; a long one means the study could not pin the effect down
Does the line cross the null? If it does, that study on its own did not reach significance

Applying this to the plot above: Ahmed 2021 sits at 0.30 with a short interval clear of the null — a precise, significant positive result. Dubois 2023 sits at 0.28, almost the same estimate, but its interval runs from −0.36 to 0.92 and crosses the null. Those two studies agree closely on the size of the effect; they differ entirely in how confidently they establish it.

A non-significant study is not a contradictory study

Dubois 2023 is frequently described as a study that “found no effect”. It found an effect of 0.28 — almost identical to the largest study in the plot — but lacked the precision to exclude zero. Treating it as evidence against the intervention misreads the plot.

Weighting

Why the squares are different sizes

Square area is proportional to the weight the study carries in the pooled estimate, and weight is determined by precision: 1 / variance. Since variance falls as sample size rises, larger studies get bigger squares.

StudyWeightWhat that means
Ahmed 202142.1%Contributes more than the bottom three combined
Bianchi 202224.3%Substantial
Costa 202014.8%Moderate
Dubois 202310.5%Limited influence
Evans 20198.3%Minimal influence

The visual purpose of the squares is to stop your eye from treating all rows equally. Without them, five studies look like five equal votes; with them, it is immediately clear that one study accounts for over 40% of the result.

Weights depend on the model

Under a fixed effect model, weights are driven purely by precision, so large studies dominate heavily. Under random effects, the between-study variance is added to every study's variance before weighting, which compresses the differences and gives small studies relatively more influence. The same five studies can produce visibly different squares depending only on the model chosen.

Convert your studies into poolable effect sizes

Our effect size calculator turns means, SDs and group sizes into Cohen's d and Hedges' g with confidence intervals — the inputs a forest plot needs.

Use the calculator

The pooled result

The diamond, and the line of no effect

The diamond represents the pooled estimate. Its centre is the point estimate and its left and right points mark the confidence interval, so a narrow diamond means a precisely estimated pooled effect.

Diamond positionInterpretation
Entirely to the right of the nullSignificant effect favouring the intervention
Entirely to the leftSignificant effect favouring the control
Touching or crossing the nullPooled effect not statistically significant

In the plot above the diamond sits at 0.39 with an interval of [0.24, 0.54], clear of the null line — a significant pooled effect of moderate size.

The diamond shape is not decorative. Because its widest point is at the centre and it tapers to the interval bounds, it draws the eye to the point estimate while still showing the range — and it is visually distinct from the study rows above, so nobody mistakes the summary for another study. Some plots additionally draw a vertical dashed line through the diamond's centre, extended up through the study rows, which makes it easy to see at a glance which individual studies sat above and below the pooled result.

Ordering matters more than it appears. Studies are most often listed alphabetically or by year, neither of which helps the reader. Sorting by effect size makes any gradient immediately visible; sorting by weight puts the influential studies at the top where they belong. Cochrane reviews conventionally use year order to show how the evidence accumulated, which is defensible — but whichever you choose, say so in the caption, because a reader who assumes one ordering and gets another will misread the pattern.

Which side is which

Always read the axis labels beneath the plot before interpreting direction. “Favours intervention” may be on the left or the right depending on how the outcome was coded, and for harmful outcomes such as mortality a lower value is the better result. Assuming right means good is a genuine and common source of error.

Some plots add a prediction interval

A second, wider line through the diamond shows the prediction interval — the range in which the effect of a future study would be expected to fall. Where heterogeneity is substantial, the diamond can be clear of the null while the prediction interval crosses it, meaning the average effect is positive but a new study might find nothing. That distinction is usually the most practically important thing on the chart.

Heterogeneity

Spotting heterogeneity by eye

Before reading the statistics, look at how the intervals line up.

Low heterogeneity

  • Intervals overlap substantially
  • Estimates cluster in a narrow band
  • All studies point the same direction
  • I² typically below 40%

High heterogeneity

  • Some intervals barely overlap
  • Estimates scattered widely
  • Studies point in opposite directions
  • I² typically above 60%

A useful habit is to look at the plot twice: once ignoring the diamond entirely, to form your own impression of what the studies show, and once with it. If your impression from the individual rows differs noticeably from the pooled result, that gap is worth understanding before citing the number. Usually it means one heavily weighted study is doing most of the work, or that a couple of imprecise studies pointing the other way are contributing almost nothing to the average despite occupying as much vertical space as the rest.

Statistics accompanying the plot usually appear beneath it as Q, and τ². Read them alongside the visual impression rather than instead of it: I² is a proportion and can be high simply because the studies were precise, so a plot of tightly estimated studies with small genuine differences may report a large I² that overstates the practical disagreement.

Judging by eye has a known weakness worth naming: the visual impression is driven by how wide the plotted axis happens to be. Software chooses that range to fit the most extreme interval, so a single very imprecise study stretches the axis and makes every other interval look narrow and tightly clustered. Two plots of the same data can therefore give quite different impressions of consistency. Check the axis values before trusting the visual impression.

When studies point in opposite directions

A pooled estimate averaging a strong positive and a strong negative finding describes neither. If the forest plot shows genuine disagreement in direction, the productive response is to investigate why — through subgroup analysis or meta-regression — rather than to report the average as though it settled the question.

Variants

Subgroup and sensitivity plots

Where an interval extends beyond the plotted axis, software draws an arrowhead rather than rescaling the whole chart around one imprecise study. That is the right visual choice, but it hides the true extent of that study's uncertainty. Always read the numeric column alongside the graphic, because a row with an arrow is telling you the study established very little.

VariantWhat it showsUse
Subgroup forest plotStudies grouped, each with its own diamondTesting whether the effect differs by a study characteristic
Cumulative forest plotStudies added one at a time, usually by yearShowing how the evidence accumulated over time
Leave-one-out plotPooled estimate with each study removed in turnTesting whether one study drives the result
Multi-outcome plotSeveral outcomes stackedSummarising a review with several endpoints

A subgroup plot shows a diamond per subgroup plus an overall diamond, and is usually accompanied by a test for subgroup differences. Note that subgroup analyses are observational even within a randomised trial — the studies were not randomly assigned to subgroups — so a difference between them is a hypothesis, not a causal finding, and it must be pre-specified to carry much weight.

Leave-one-out plots deserve wider use. If removing a single study moves the pooled estimate substantially or changes its significance, the conclusion rests on that study and the review should say so plainly.

Get your forest plot produced to publication standard

Send your extraction table. A named statistician runs the meta-analysis and returns forest, funnel and sensitivity plots with the heterogeneity properly assessed.

See statistical consultancy

Pitfalls

Six misreadings to avoid

1. Treating every row as an equal vote

Counting how many studies were significant ignores the weights entirely. Look at the square sizes.

2. Reading a wide interval as a contradictory finding

An imprecise study that agrees on the estimate but crosses the null is not evidence against the effect.

3. Assuming right means favourable

Read the axis labels. For outcomes such as mortality or relapse, a result on the left is the better one.

4. Ignoring the scale on ratio plots

Odds and risk ratios are plotted on a log axis with the null at 1. Distances are not linear.

5. Reporting the diamond without the heterogeneity

A pooled estimate from studies that disagree needs that disagreement reported alongside it.

6. Overlooking a dominant study

If one study carries 60% of the weight, the meta-analysis is largely reporting that study. Run a leave-one-out analysis and say so.

Reporting

Producing and reporting a forest plot

Blank rows and section headings earn their space in longer plots. Grouping studies by design, setting or risk of bias with a labelled heading between the groups costs nothing and reveals structure that would otherwise be buried in an undifferentiated list of thirty rows. Most meta-analysis packages support this directly.

Label both axis directions so the reader knows which side favours what
Include the weight column — without it the squares are hard to quantify
Show numeric estimates and intervals as text alongside the plot
State the model, fixed effect or random effects, in the caption
Report Q, I² and τ² beneath the plot
Give k and total N in the caption
Add a prediction interval where heterogeneity is substantial

One practical caution on producing your own: the plot must be generated from the same analysis that produced your reported numbers, not assembled separately. Hand-drawn or spreadsheet-built plots drift out of step with the analysis as soon as a study is added or a model changed, and a figure that disagrees with the text is the kind of inconsistency reviewers notice immediately.

Standard software includes the metafor and meta packages in R, the metan suite in Stata, and RevMan for Cochrane reviews. All produce publication-standard plots directly; drawing one by hand in a graphics program is both laborious and prone to introducing errors of scale.

A caption that does the work

“Forest plot of five studies (N = 1,758) comparing retrieval practice with conventional revision on end-of-module performance. Random effects model. Squares are proportional to study weight; the diamond shows the pooled estimate. Pooled SMD = 0.39, 95% CI [0.24, 0.54]; Q(4) = 3.12, p = .538, I² = 0%.”

Answers

Frequently asked questions

What is a forest plot used for?

To display the results of every study in a meta-analysis on a single scale alongside the pooled estimate. It lets a reader see each study's effect and precision, how much weight each carries, and whether the studies agree — information a pooled number alone cannot convey.

What does the diamond on a forest plot mean?

It represents the pooled estimate from the meta-analysis. Its centre is the point estimate and its left and right points mark the confidence interval. If the diamond touches or crosses the line of no effect, the pooled result is not statistically significant.

Why are the squares different sizes?

Square area is proportional to the weight each study carries, which is determined by its precision — the inverse of its variance. Larger, more precise studies get bigger squares and influence the pooled estimate more. The sizing exists to stop readers treating every study as an equal vote.

What does it mean when a confidence interval crosses the line of no effect?

That study did not reach statistical significance on its own. It does not mean the study found no effect — the point estimate may be substantial — only that the interval is wide enough to include the possibility of no effect, usually because the study was small.

How do I tell if there is heterogeneity from a forest plot?

Look at whether the confidence intervals overlap. Substantial overlap with estimates clustered together suggests low heterogeneity; scattered estimates with barely overlapping intervals, especially pointing in opposite directions, suggest high heterogeneity. Confirm against the I squared and tau squared reported beneath the plot.

Why is the axis logarithmic on some forest plots?

Because ratio measures such as odds ratios and risk ratios are not symmetric on a linear scale — halving and doubling should appear as equal distances from the null. A log axis achieves that, and the null line sits at 1 rather than 0.

What software makes forest plots?

The metafor and meta packages in R, the metan suite in Stata, and RevMan for Cochrane reviews, all of which produce publication-standard output directly. Drawing one manually in a graphics program risks introducing errors of scale and is not recommended.

What is the difference between a confidence interval and a prediction interval on a forest plot?

The confidence interval, shown by the diamond's width, describes how precisely the average effect has been estimated. The prediction interval, sometimes drawn as a wider line through the diamond, describes where the true effect in a new setting would be expected to fall. Where heterogeneity is substantial the prediction interval is much wider and is usually the more useful figure.

Send the data. Get a fixed quote.

Attach your dataset or just describe the project. A named statistician replies with a price and a deadline, usually within one working day.