Systematic review / Meta-analysis

How to Do a Meta-Analysis: Searching and Screening, Pooling, Heterogeneity, Publication Bias and Forest Plots

This page is for medical graduate students doing their first systematic review and meta-analysis. It sets out the decision criteria for each step according to the Cochrane Handbook version 6.5, PRISMA 2020 and GRADE, and compares the defaults of RevMan, Stata and R. The R part was run with meta 8.5-0 and metafor 5.0.1 on public example data.

Short answer

A meta-analysis proceeds as follows: register the protocol in PROSPERO (at the latest before data extraction starts), search PubMed, Embase, CENTRAL and Chinese databases and report the full search strategies, screen and extract in duplicate, assess risk of bias with RoB 2 or ROBINS-I, choose the effect measure (RR or OR for binary outcomes, MD or SMD for continuous outcomes, HR for survival data), then pool. Use a random-effects model with REML for tau², the Hartung-Knapp confidence interval when tau²>0 and there are more than two studies, and report a prediction interval when there are five or more studies. Judge heterogeneity by tau² and the prediction interval; I² only tells you what share of the observed variation comes from between-study differences. Funnel plot tests and meta-regression need at least 10 studies. Finally, rate the certainty of evidence for each outcome with GRADE.

Workflow

The output of each step has to be reported under the 27 items of PRISMA 2020. The 'common errors' column draws on the Cochrane Handbook, the PRISMA 2020 explanation paper and PROSPERO guidance.

StepWhat to doDecision criteria and common errorsBasis
1 Register the protocolEnter PICO, sources, eligibility criteria, primary outcomes, pooling methods, and planned subgroup and sensitivity analyses in PROSPEROPROSPERO requires submission before data extraction starts; subgroup analyses not in the protocol must be labelled post hoc in the paper. The new system launched in March 2025 publishes records automatically once all authors approve them and flags similar topics at submissionPROSPERO guidance; Handbook 10.11.5
2 SearchPubMed/MEDLINE, Embase and Cochrane CENTRAL, plus Chinese databases (CNKI, Wanfang, VIP, SinoMed) and trial registries (ClinicalTrials.gov, ChiCTR)PRISMA 2020 item 7 requires the full search strategy for every database, including limits; the 2009 version asked for one database only. Give search dates to the dayPRISMA 2020 item 7
3 DeduplicateImport into EndNote, Zotero, Rayyan or Covidence and remove duplicatesMerge the conference abstract, main report and follow-up reports of one trial into a single 'study' and count by study to avoid double countingHandbook chapter 5
4 ScreenTitle and abstract screening, then full text, by two reviewers independently, with disagreements resolved by discussion or a third reviewerPRISMA 2020 item 8 asks how many reviewers screened each record, whether independently, and whether automation tools were used; the flow diagram lists full-text exclusions with reasons and countsPRISMA 2020 items 8, 16a
5 Extract dataDesign the extraction form in advance, pilot it on 2–3 papers, then extract in duplicateWhen an outcome is reported both adjusted and unadjusted, follow the priority set in the protocol; for multi-arm trials decide first which arms are relevantHandbook chapters 5, 6, 23
6 Risk of biasRoB 2 for RCTs, ROBINS-I for non-randomised studies of interventions (V2 released in 2025), ROBINS-E for exposures, QUADAS-2 for diagnostic studiesAssess per outcome: an RoB 2 judgement refers to a specific result for a specific outcome, not to the whole paperSterne 2019 BMJ; Cochrane ROBINS-I V2
7 Pooling and heterogeneityChoose the effect measure; random effects + REML + HK; report tau², I² and the prediction intervalDo not choose between fixed and random effects based on the heterogeneity testHandbook 10.10.4
8 Publication bias, sensitivity analysisFunnel plot and tests with ≥10 studies; sensitivity analyses as plannedUse trim-and-fill only as a sensitivity analysis, not as the adjusted main resultHandbook 13.3; Peters 2007
9 Certainty of evidenceRate each outcome with GRADE and build a summary of findings (SoF) tablePRISMA 2020 added items 15 and 22 on the methods and results of certainty assessmentPRISMA 2020; Core GRADE 2025

Risk of bias

Study typeRecommended toolHow it judgesWhat to know
Randomised controlled trialsRoB 2 (2019)Five domains: randomisation process, deviations from intended interventions, missing outcome data, outcome measurement, selective reporting; overall low risk, some concerns or high riskThe older Cochrane risk of bias tool (seven items, 2011) is still common in Chinese papers and reviewers may ask for RoB 2; RoB 2 first requires deciding whether the effect of interest is assignment (ITT) or adherence
Non-randomised studies of interventions (cohorts or case-control studies comparing treatments)ROBINS-I (2016); V2 released in 2025Uses a target trial as reference and assesses seven domains including confounding, selection and classification of interventions; V2 uses algorithms to suggest judgements from signalling questions and covers immortal time biasROBINS-I takes longer than NOS but distinguishes whether confounding was adequately handled
Exposure studies (for example smoking and disease)ROBINS-E (2024)Similar structure to ROBINS-I, for exposuresThe tool is new and Chinese examples are few
Observational studies (a common choice in Chinese meta-analyses)Newcastle-Ottawa Scale (NOS)Selection, comparability and outcome, up to nine starsStang 2010 found that NOS items and scoring lack justification and inter-rater agreement is low; report the judgement for each item rather than only the star total, and do not exclude studies by a '≥7 stars = high quality' rule
Diagnostic accuracy studiesQUADAS-2Risk of bias and applicability in four domainsDiagnostic meta-analysis uses different pooling methods (bivariate models) from the intervention meta-analysis on this page

Effect measures

Data typeOptionsHow to chooseBasis
BinaryRR, OR, RDUse RR or OR by default. Relative effects are more consistent across baseline risks than absolute effects, and pooled RDs tend to be more heterogeneous; the Handbook advises against pooling RDs unless there is reason to expect them to be consistent. An OR cannot be read as an RR, and the two diverge noticeably when event rates exceed 10%. Absolute effects in the SoF table are obtained by applying the RR to a baseline riskHandbook 10.4.3
Continuous, same unitMD (mean difference)Use MD for the same scale and unit; results read directly in clinical units. Change scores and final values can be pooled togetherHandbook 10.5.2
Continuous, different scalesSMD (Hedges g)Use SMD only when studies measure the same construct on different scales. The SMD default in RevMan, meta and Stata is Hedges g (with small-sample correction). Change scores and final values cannot be mixed in an SMD because their SDs mean different things. When scales run in opposite directions, multiply one side by -1 firstHandbook 6.5.1.2, 10.5.2
SurvivalHRPool ln(HR) and its SE (generic inverse variance). Do not replace HR with survival at one time point unless all studies report only that time point. Data sources in order of preference are in the next sectionHandbook 6.8.2; Tierney 2007
Counts and ratesRate ratioTime units cancel out in a rate ratio, so harmonising them does not change the resultHandbook 6.7.1

Data conversion

Ask the corresponding author for the data first. If that fails, convert as below, and remove all converted studies once in a sensitivity analysis to see whether the conclusion changes.

The team behind the Wan, Luo and Shi methods (Tong Tiejun and colleagues, Hong Kong Baptist University) provides an online calculator and Excel sheet that test for skewness before converting; R meta (the median, q1, q3, min and max arguments of metacont) and metafor (conv.fivenum) implement the same formulae.

Handbook 10.5.3 gives a skewness check: for measures that can only be positive, a mean/SD ratio below 2 suggests skew and below 1 indicates clear skew; use the reported mean and SD of such studies in an MD analysis with caution, and do not mix log-transformed and untransformed data in one meta-analysis.

Harmonise units before pooling. For example, divide glucose in mg/dL by 18 to get mmol/L, and total cholesterol in mg/dL by 38.67. Pooling with SMD does not replace unit conversion: SMD treats SD differences caused by different units as differences in effect.

Entering medians and quartiles directly in metar
# With only medians and quartiles, enter them directly in metacont; meta converts internally
# Defaults: mean by Luo 2018; SD by Shi 2020 (when the range is available) or Wan 2014 (IQR only)
d <- data.frame(study = c("A", "B", "C"),
                n.e = c(40, 60, 35), med.e = c(12, 15, 11), q1.e = c(8, 10, 7), q3.e = c(20, 24, 18),
                n.c = c(42, 58, 33), med.c = c(16, 19, 14), q1.c = c(11, 13, 9), q3.c = c(25, 30, 22))
mq <- metacont(n.e, median.e = med.e, q1.e = q1.e, q3.e = q3.e,
               n.c = n.c, median.c = med.c, q1.c = q1.c, q3.c = q3.c,
               data = d, studlab = study, sm = "MD")
cbind(mq$mean.e, mq$sd.e)          # study A: 13.42, 9.21
# For skewed data, switch to the McGrath 2020 Box-Cox method (requires estmeansd)
update(mq, method.mean = "BC-McGrath", method.sd = "BC-McGrath")
# Equivalent metafor function
conv.fivenum(q1 = q1.e, med = med.e, q3 = q3.e, n = n.e, data = d)
What the paper reportsMethodAssumptions and limits
Median + quartiles (Q1, Q3)Mean: Luo 2018; SD: Wan 2014 (IQR divided by a factor that depends on sample size and approaches 1.35 for large n)These methods assume approximate normality. Skewed data lead to underestimated means and SDs; test skewness with Shi 2023 and, if skew is clear, use the McGrath 2020 Box-Cox or quantile estimation methods
Median + minimum and maximumMean: Luo 2018; SD: Wan 2014Handbook 6.5.2.6 recommends not estimating SDs from ranges, which are driven by extreme values
Median + quartiles + range (five-number summary)Mean: Luo 2018; SD: Shi 2020Uses the most information and gives the smallest error
Mean + 95% CISD = √n × (upper − lower) / (2 × t₀.₉₇₅,ₙ₋₁); for n≥60 use 1.96 for tValid only when the CI was computed from the t distribution
Two group means + between-group P or t valueBack-calculate the between-group SE from t, then the pooled SD of both groupsWhen a paper only says P<0.05, using P=0.05 overestimates the SD and makes the result conservative
Only subgroup n, means and SDsCombine into one group with the formulae in Handbook table 6.5.aDo not include the two subgroups as two separate studies
Cox regression HR with 95% CIUse ln(HR) and SE = (ln upper − ln lower)/3.92 directlyFirst choice
Only log-rank P value and number of eventsTierney 2007: variance of ln(HR) ≈ 4/total events (1:1 allocation)Requires the direction of effect and the allocation ratio
Only Kaplan-Meier curvesDigitise with WebPlotDigitizer or Engauge Digitizer, then use the Tierney spreadsheet (Parmar method) or the Guyot 2012 algorithm to reconstruct individual data and fit a Cox model; the R package IPDfromKM implements the reconstruction after digitisingGuyot 2012 validation shows that HRs are estimated reasonably only when the numbers at risk or total events are reported; without both, the reconstructed data assume no censoring, events are overestimated and the CI is too narrow

Pooling model

Handbook section 10.10.4 is explicit: the choice between fixed and random effects must not be based on a statistical test for heterogeneity. 'Use random effects if I²>50%, otherwise fixed effects' is the most common wrong practice in Chinese tutorials.

DecisionRecommendationBasis and conditions
Fixed or randomReport random effects by default; state this in the protocol and run fixed effects as a sensitivity analysisHandbook 10.10.4.1 gives no universal recommendation but rejects choosing by heterogeneity test. The two models weight studies differently: random effects give small studies more weight, so the results separate when small-study effects exist
tau² estimatorREMLSince the 2024 update the Handbook uses REML as the default (10.10.4.4), and RevMan has defaulted to REML since January 2025. Langan 2019 compared nine estimators and recommended REML; Veroniki 2016 found Paule-Mandel performs well for binary and continuous data. DL underestimates tau² with many small studies or rare events
Confidence intervalHartung-Knapp (HKSJ, t distribution with k−1 df) when tau²>0 and there are more than two studiesIn IntHout 2014 simulations the type I error of DL could exceed 30% while HKSJ at most doubled the nominal rate; among 689 Cochrane meta-analyses, 25.1% of results significant with DL were not significant with HKSJ
Only two studiesDo not use HK; use a Wald-type interval and compare methods in a sensitivity analysisWith k=2 the t distribution has 1 df and the interval is extremely wide (tested on this page: RR CI 0.004–19.2)
tau²=0 or Q < k−1Use the Wald-type interval or the modified HK (mKH)In this situation the HK interval can be narrower than Wald (tested on this page). Röver 2015 recommends the modified version with few studies of varying precision
Prediction intervalReport when there are ≥5 studies and no clear funnel plot asymmetryHandbook 10.10.4.3; use k−1 df (previously k−2). The prediction interval answers where the true effect of a new similar study is likely to fall
CI for tau²Q-profile methodInformative only with ≥5 studies
With fewer than five studies of very different sizes, interpret results cautiously even with HKSJ (IntHout 2014).

Software differences

Most mismatches come from different defaults. In the methods section, state the tau² estimator, the CI method, the distribution and df of the prediction interval, and how I² was computed.

Tested on this page (dat.bcg, 13 studies, REML): the default prediction interval in meta is RR 0.136–1.762 (t distribution, k−1); with Stata's K−2 df it is 0.134–1.785; metafor with test = "z" gives 0.155–1.549. For the same data, I² computed from Q in meta is 92.1% (95% CI 88.3%–94.7%), and computed from tau² in metafor is 92.2% (CI 81.9%–97.7%).

Problems when copying old tutorials: from meta 6.0, hakn = TRUE became method.random.ci = "HK"; comb.fixed became fixed in 5.0 and common in 5.5; byvar became subgroup in 5.0. In 8.5 the old arguments still run but give a deprecated warning; the file argument of forest() was renamed filename in 8.5.

SettingRevMan (from January 2025)Stata 16+ metaR meta 8.5R metafor 5.0
Default tau² estimatorREML (old RevMan 5 used DL)REMLREMLREML
Random-effects CIPrompts HKSJ when tau²>0 and k>2Wald; add se(khartung) for HKWald (method.random.ci = "classic"); change to "HK"z test; test = "knha" for HK
Prediction intervalt distribution, k−1predinterval option, t distribution, K−2prediction = TRUE, t distribution, k−1 (from 8.0; k−2 before)predict(): normal distribution with test = "z", t distribution with test = "knha"
I² formulaFrom tau² (new)From tau² in random-effects modelsFrom Q by default; method.I2 = "tau2" to changeFrom tau²
Zero eventsMH, Peto; 0.5 continuity correctionmhaenszel, peto; correction value configurableMH without correction by default; Peto and GLMM availablerma.mh, rma.peto, rma.glmm
Egger test / trim-and-fillNot availablemeta bias, egger; meta trimfillmetabias; trimfillregtest; trimfill
Meta-regressionNot availablemeta regressmetaregrma(mods = ~ …)

Heterogeneity

Subgroup analysis: specify subgroup variables in the protocol; judge by the test for subgroup differences and do not compare P values within subgroups (one significant and one non-significant subgroup does not show a difference). With very few studies per subgroup, within-subgroup random effects can hardly be estimated. In this page's test, splitting by allocation method into three groups gave P<0.0001 for subgroup differences under the common-effect model and P=0.39 under the random-effects model, opposite conclusions; the subgroup with two studies had an HK CI of 0.016–20.8.

Meta-regression: generally not done with fewer than 10 studies, and roughly 10 studies are needed per covariate (Handbook 10.11). The relationship between a study-level covariate (such as mean age) and the outcome may not reflect the individual-level relationship, which is ecological bias. Use the Knapp-Hartung adjustment for meta-regression P values (test = "knha" in metafor, se(khartung) in Stata).

Core GRADE (2025) judges inconsistency by the spread of point estimates, the overlap of confidence intervals and on which side of the decision threshold each study falls, not by I² alone.

MeasureQuestion it answersCommon misreading
Q testWhether observed differences exceed sampling errorWith few studies the test has low power, so P>0.10 does not mean no heterogeneity; with many large studies even small differences become significant
I²What share of the total observed variation comes from true between-study differencesI² is relative and depends on study precision: with tau² unchanged, I² rises towards 100% as study sizes grow (Rücker 2008). I²=50% cannot tell you whether effects range from 40 to 60 or from 10 to 90 (Borenstein 2017). The Handbook bands (0–40%, 30–60%, 50–90%, 75–100%) overlap and are only a rough guide
tau², tauVariance and SD of true effects across studies, on the effect-size scaletau² estimates are unstable with few studies; report the Q-profile CI
Prediction intervalRange of the distribution of true effectsA significant pooled effect with a prediction interval crossing the null means the effect may be absent in some settings; this belongs in the discussion

Publication bias

Tested on this page (dat.bcg, k=13): Egger P=0.19, Harbord P=0.32, Peters P=0.23, none significant. Trim-and-fill with meta defaults (ma.common = TRUE, iterating with a common-effect model) imputed 4 studies and moved the RR from 0.49 (0.34–0.70) to 0.71 (0.46–1.10), no longer significant; with ma.common = FALSE or metafor's trimfill on an rma object it imputed 1 study, RR 0.52 (0.37–0.74). The heterogeneity in these data comes mainly from latitude (see meta-regression), which is the situation described by Peters 2007.

MethodWhen to useLimitations
Funnel plotAt least 10 studies; not when studies are of similar sizeAsymmetry shows small-study effects, which may come from publication bias, poorer quality of small studies, true heterogeneity, artefactual correlation between effect size and SE, or chance (Handbook table 13.3.b). Contour-enhanced funnel plots show whether the missing region lies in the non-significant zone
Egger regression testMD for continuous outcomes; ≥10 studiesFor OR and SMD the effect size is correlated with its SE, so the original Egger test gives false positives; power is low with few studies
Harbord / Peters testsBinary ORPeters uses 1/total sample size as the predictor, avoiding the artefactual correlation with SE
Trim-and-fillSensitivity analysis onlyAssumes asymmetry is entirely due to publication bias; with substantial heterogeneity it may impute studies and underestimate the effect even without publication bias (Peters 2007). Defaults differ between software and results can differ greatly (see this page's test)

Sensitivity analysis and GRADE

Certainty ends up as high, moderate, low or very low; build the summary of findings table with GRADEpro GDT.

GRADE factorDirectionHow to judge
Starting point—RCTs start at high, observational studies at low (when ROBINS-I is used they may start at high and be rated down for bias)
Risk of biasDownWhether heavily weighted studies are at high risk
InconsistencyDownPoint estimates spread and CIs do not overlap, not explained by prespecified subgroups
IndirectnessDownPopulation, intervention, comparator or outcome differ from the question
ImprecisionDownThe 95% CI crosses the minimal important difference (MID) or the null (Core GRADE 2)
Publication biasDownAsymmetric funnel plot, or only small industry-funded studies
Large effect, dose-response, confounding that would reduce the effectUpMainly for observational studies
  • Change the model: fixed and random effects; tau² by REML, DL and PM.
  • Remove studies at high risk of bias (high risk in RoB 2, serious or critical in ROBINS-I).
  • Remove studies with converted data (means estimated from medians, HRs reconstructed from curves).
  • Handle zero-event studies with another method (see common errors).
  • Leave-one-out: remove each study in turn and report the range of pooled effects and whether any single study changes the conclusion; use it to find influential studies, not to pick the subset that lowers I².
  • For the primary outcome, analyse adjusted effects only and unadjusted effects only.

Common errors

Double counting a shared control group in multi-arm trials

When doses A and B are both compared with the same control and both comparisons enter one meta-analysis, the control participants are counted twice, their weight is inflated and the two results are correlated. Handbook 23.3.4 recommends combining the relevant experimental groups into one; splitting the control group only partly solves the problem.

Including several reports of the same population

Main reports, subgroup reports and long-term follow-ups often have different author orders and titles. Check registration numbers, recruitment periods and sample sizes, and include each trial once.

Mixing adjusted and unadjusted effects

In observational studies, some report multivariable-adjusted HRs and others only crude HRs or raw counts; pooling them treats results with different confounding control as one effect. Specify in the protocol that the most fully adjusted result is preferred, and run a sensitivity analysis with adjusted results only.

Mishandling zero events

With zero events in one arm, inverse variance methods need a continuity correction, and a 0.5 correction pulls results towards the null. Studies with zero events in both arms are dropped by default for OR and RR. With event rates below 1% and similar group sizes the Peto method has the least bias; with unbalanced groups or large effects use MH without correction or a GLMM (Bradburn 2007). In this page's test on the rosiglitazone data, the choice of method changes the conclusion.

Taking SE for SD

Using 'mean ± SE' from a table as the SD underestimates the SD by a factor of √n and inflates that study's weight and SMD. SD = SE × √n. Check table footnotes during extraction.

Inconsistent units and directions

MDs in different units cannot be pooled directly; if scale directions (higher is better versus higher is worse) are not aligned, the pooled effect cancels out.

Removing studies to lower I²

Deleting studies one by one until I² falls below 50% and reporting the rest is outcome-driven selection. Explain heterogeneity with prespecified subgroups or meta-regression; what cannot be explained goes into the discussion and leads to rating down in GRADE.

Funnel plots and meta-regression with too few studies

An Egger test on 5–6 studies with P>0.05 is reported as 'no publication bias'. With fewer than 10 studies, state that the test was not done and why.

R code

The code runs as is; the example data install with the metadat package (install.packages(c("meta", "metafor"))). For your own binary outcome data you need four columns: events and totals for each group.

meta_bcg.Rr
# R 4.5.3 + meta 8.5-0 + metafor 5.0.1; example data dat.bcg from the metadat package (13 BCG vaccine trials)
library(meta); library(metafor)
data(dat.bcg, package = "metadat")
bcg <- dat.bcg; bcg$study <- paste(bcg$author, bcg$year)

# 1. Pooling: REML for tau2, both Wald and Hartung-Knapp CIs, plus a prediction interval
m <- metabin(tpos, tpos + tneg, cpos, cpos + cneg, data = bcg, studlab = study,
             sm = "RR", method = "Inverse", method.tau = "REML",
             method.random.ci = "HK", prediction = TRUE)
summary(m)
update(m, method.random.ci = "classic")   # Wald-type interval for comparison

# 2. Sensitivity analysis across tau2 estimators
for (mt in c("DL", "REML", "PM")) print(update(m, method.tau = mt, method.random.ci = "HK"))

# 3. Forest plot (from meta 8.5 the file argument is called filename)
forest(m, sortvar = TE, prediction = TRUE, print.tau2 = TRUE,
       filename = "forest_bcg.png")

# 4. Funnel plot and small-study tests (only with k >= 10; use Harbord or Peters for OR outcomes)
funnel(m, type = "contour", common = FALSE)
metabias(m, method.bias = "Egger")
metabias(m, method.bias = "Peters")

# 5. Trim-and-fill: state the model used in the iterations, otherwise meta and metafor disagree
trimfill(m, ma.common = FALSE)          # random-effects iterations, same as metafor
trimfill(m)                              # meta default: ma.common = TRUE (common-effect iterations)

# 6. Subgroups and meta-regression (meta-regression only with >= 10 studies)
update(m, subgroup = alloc)
metareg(m, ~ ablat)

# 7. Leave-one-out
metainf(m, pooled = "random")

# 8. metafor syntax: Hartung-Knapp via test = "knha"; the prediction interval then uses the t distribution
es <- escalc(measure = "RR", ai = tpos, bi = tneg, ci = cpos, di = cneg, data = bcg)
r <- rma(yi, vi, data = es, method = "REML", test = "knha")
predict(r, transf = exp)
regtest(r); trimfill(r); leave1out(r); confint(r)
# If REML fails (Fisher scoring algorithm did not converge):
# rma(yi, vi, data = es, control = list(stepadj = 0.5, maxiter = 1000))

Tested on this page

Conditions: R 4.5.3, meta 8.5-0, metafor 5.0.1, metadat 1.6-0, IPDfromKM 0.1.10, macOS arm64. Data are dat.bcg (13 BCG vaccine trials, RR) and dat.nissen2007 (42 rosiglitazone trials, myocardial infarction) shipped with metadat.

ComparisonResultNote
tau²: DL / REML / PM / ML0.309 / 0.313 / 0.318 / 0.280Pooled RR is about 0.49 in all cases; with 13 studies the estimator has little effect on the point estimate
Wald vs HK (REML)RR 0.489, CI 0.344–0.696 (P=7×10⁻⁵) vs 0.330–0.726 (P=0.002)HK widens the CI and increases P about 27-fold
Weights under fixed and random effectsLargest study TPT Madras: 41.4% fixed, 10.2% randomFixed-effect RR 0.65, random-effects 0.49; the gap comes from stronger effects in small studies
Two studiesWald 0.134–0.527; HK 0.004–19.2Do not use HK with k=2
Four studies, tau²=0, Q=1.04<3Wald 0.188–0.310; HK 0.190–0.307; modified HK 0.160–0.363HK is narrower than Wald; the modified version corrects this
Meta-regression (latitude)ln(RR) falls 0.029 per degree of latitude; R²=75.6%, residual I²=68%With KNHA adjustment P goes from <0.0001 to 0.005
Leave-one-outRR 0.452 (without TPT Madras) to 0.533 (without Hart & Sutherland)No single study changes the conclusion
Zero events (42 trials, 4 double-zero, 26 single-zero)MH OR 1.28 (0.95–1.73); Peto 1.43 (1.03–1.98); IV + 0.5 correction 1.29 (0.94–1.76); GLMM 1.43 (1.03–1.98)Peto and GLMM are significant, MH and IV are not; several trials in these data used 2:1 allocation, so the balance assumption of Peto is not met; report several methods side by side
SMD: Hedges g vs Cohen d−0.536 vs −0.544 (9 studies)Differences arise mainly in small studies with n<30
Median to mean (lognormal, n=200, true mean 23.3, SD 15.8)Wan/Luo: 20.4, 11.5; Box-Cox (McGrath): 22.5, 13.3With skewed data the normal-based methods underestimate the SD by about 27%
HR reconstructed from KM curves (simulated data, true HR 0.568)With risk table: 0.572 (0.417–0.785), events 160/163; without: 0.565 (0.435–0.735), events 231/163Without numbers at risk the events are overestimated by 42% and the CI is too narrow, consistent with Guyot 2012

Chinese community

The items below come from a Chinese methods journal article, CSDN columns and tools provided by method authors, and only parts that agree with the Cochrane Handbook or the original methods papers are included. The many templated 2025–2026 'meta-analysis pitfall guides' on CSDN were not used.

'Deep extraction' of continuous data

Liu Haining and colleagues (Chinese Journal of Evidence-Based Medicine, 2017, vol. 17, no. 1) cover seven situations: estimating from the median with extremes or quartiles, reading values from figures, combining subgroups, back-calculating the SD from a 95% CI of the mean, and back-calculating a pooled SD from a between-group P or t value, with Stata and Excel commands for t quantiles (invttail, T.INV.2T). They note that using 0.05 when only P<0.05 is reported overestimates the SD. This agrees with Handbook chapter 6.

Extracting HRs with the Tierney spreadsheet

A CSDN column (February 2024) gives the steps after digitising with Engauge Digitizer: enter coordinates and follow-up in Sheet 2a, compare the simulated and original curves in Sheet 2b and re-digitise if they differ, and read the HR and CI in Sheet 4. This agrees with the instructions in the Tierney paper.

Tong Tiejun's online calculator tests skewness first

The online calculator from the authors of Wan 2014, Luo 2018 and Shi 2020 (Zeng Xiantao is also a Shi 2020 author) asks you to click 'Detect the skewness' first, using the Shi 2023 method, and warns against normal-based methods when skew is significant. Many Chinese tutorials quote the formulae and skip this step.

'Use random effects when I²>50%' and 'nine studies are enough for a funnel plot'

Several CSDN and WeChat tutorials repeat these two rules. Handbook 10.10.4.1 explicitly rejects choosing the model by heterogeneity test, and chapter 13 requires at least 10 studies for funnel plot tests. Reviewers often point out both.

Forest plot weights do not depend only on sample size

The Zhihu article ranked first on Bing for 'forest plot' (森林图) says 'the larger the sample, the larger the weight'. This holds only approximately under the fixed-effect model; the random-effects model adds tau² to each study's variance and pulls down the weight of large studies (in this page's test the largest study fell from 41.4% to 10.2%). When reading a forest plot, check which model each weight column belongs to.

Hand to an agent

Search strategy drafting, data conversion, pooling and plotting can be handed to the Scientify scientific agent; inclusion decisions and risk of bias judgements still need you and your co-reviewers.

Example instruction: “Following the attached PROSPERO protocol, run the meta-analysis on the extraction sheet of 12 RCTs: RR for the primary outcome, random effects + REML + Hartung-Knapp, with a prediction interval; convert medians and quartiles with Luo/Wan after a skewness test; run the planned subgroups, leave-one-out and a sensitivity analysis without high risk of bias studies; output forest plots, funnel plots and a draft methods section.”

The agent installs R meta, metafor and other packages on an isolated cloud computer, reads the extraction sheet, checks for SE/SD confusion, inconsistent units and double-counted multi-arm trials, completes conversion, pooling and all sensitivity analyses, and outputs figures and a methods section mapped to PRISMA 2020 items. The agent reviews its results adversarially, for example checking whether REML and DL, HK and Wald agree, and whether the trim-and-fill iteration model is stated. The workspace keeps scripts, parameters, logs and data for reproduction, and the task keeps running after you close your computer.

You still need to check: inclusion and exclusion decisions, risk of bias judgements, whether the direction and extracted values of effects match the papers, and the reasons for GRADE ratings.

Sources

FAQ

Can I still pool when I² is high?

Whether to pool depends on whether studies are similar enough in population, intervention and outcome, not on I². If pooling is clinically sensible, use a random-effects model, report tau² and the prediction interval, explain heterogeneity with prespecified subgroups or meta-regression, and rate down for inconsistency in GRADE if it remains unexplained. When effects point in opposite directions and differ greatly, use a narrative synthesis only.

I have only three studies. What method should I use?

Use random effects with REML and give both Wald-type and Hartung-Knapp CIs as a sensitivity analysis; tau² is very unstable, so do not run funnel plots, Egger tests or meta-regression, and do not report a prediction interval (the Handbook suggests ≥5 studies). With only two studies, do not use HK.

Does an Egger test with P>0.05 show there is no publication bias?

No. With fewer than 10 studies the test has low power, and asymmetry may come from small-study quality or heterogeneity anyway. Write 'no evidence of small-study effects was found' and state the number of studies and the method (Harbord or Peters for OR outcomes).

Why do RevMan, Stata and R give different results?

Common reasons are the tau² estimator (old RevMan 5 uses DL, the others default to REML), whether HK is used, the prediction interval df (k−1 or k−2), the I² formula and the zero-event correction. Once these settings are aligned, results agree to two or three decimal places.

A paper gives only survival curves and no HR. What can I do?

Digitise with WebPlotDigitizer or Engauge Digitizer, then reconstruct individual data with the Tierney 2007 spreadsheet or IPDfromKM and compute the HR. Always enter the numbers-at-risk table below the curve and the total events; without both, the HR CI is too narrow. Remove all reconstructed studies once in a sensitivity analysis.

What if I registered in PROSPERO too late?

PROSPERO only accepts protocols submitted before data extraction starts. If extraction has begun, make the protocol public on a platform such as OSF and report the registration date and protocol changes honestly in the paper (PRISMA 2020 item 24).

Hand the computational part of your meta-analysis to Scientify

The scientific agent runs data conversion, pooling, sensitivity analyses and plotting on an isolated cloud computer, keeps all scripts, parameters and logs, and reviews the results adversarially. New users get US$5 of free credit.