Workflow
The output of each step has to be reported under the 27 items of PRISMA 2020. The 'common errors' column draws on the Cochrane Handbook, the PRISMA 2020 explanation paper and PROSPERO guidance.
| Step | What to do | Decision criteria and common errors | Basis |
|---|---|---|---|
| 1 Register the protocol | Enter PICO, sources, eligibility criteria, primary outcomes, pooling methods, and planned subgroup and sensitivity analyses in PROSPERO | PROSPERO requires submission before data extraction starts; subgroup analyses not in the protocol must be labelled post hoc in the paper. The new system launched in March 2025 publishes records automatically once all authors approve them and flags similar topics at submission | PROSPERO guidance; Handbook 10.11.5 |
| 2 Search | PubMed/MEDLINE, Embase and Cochrane CENTRAL, plus Chinese databases (CNKI, Wanfang, VIP, SinoMed) and trial registries (ClinicalTrials.gov, ChiCTR) | PRISMA 2020 item 7 requires the full search strategy for every database, including limits; the 2009 version asked for one database only. Give search dates to the day | PRISMA 2020 item 7 |
| 3 Deduplicate | Import into EndNote, Zotero, Rayyan or Covidence and remove duplicates | Merge the conference abstract, main report and follow-up reports of one trial into a single 'study' and count by study to avoid double counting | Handbook chapter 5 |
| 4 Screen | Title and abstract screening, then full text, by two reviewers independently, with disagreements resolved by discussion or a third reviewer | PRISMA 2020 item 8 asks how many reviewers screened each record, whether independently, and whether automation tools were used; the flow diagram lists full-text exclusions with reasons and counts | PRISMA 2020 items 8, 16a |
| 5 Extract data | Design the extraction form in advance, pilot it on 2–3 papers, then extract in duplicate | When an outcome is reported both adjusted and unadjusted, follow the priority set in the protocol; for multi-arm trials decide first which arms are relevant | Handbook chapters 5, 6, 23 |
| 6 Risk of bias | RoB 2 for RCTs, ROBINS-I for non-randomised studies of interventions (V2 released in 2025), ROBINS-E for exposures, QUADAS-2 for diagnostic studies | Assess per outcome: an RoB 2 judgement refers to a specific result for a specific outcome, not to the whole paper | Sterne 2019 BMJ; Cochrane ROBINS-I V2 |
| 7 Pooling and heterogeneity | Choose the effect measure; random effects + REML + HK; report tau², I² and the prediction interval | Do not choose between fixed and random effects based on the heterogeneity test | Handbook 10.10.4 |
| 8 Publication bias, sensitivity analysis | Funnel plot and tests with ≥10 studies; sensitivity analyses as planned | Use trim-and-fill only as a sensitivity analysis, not as the adjusted main result | Handbook 13.3; Peters 2007 |
| 9 Certainty of evidence | Rate each outcome with GRADE and build a summary of findings (SoF) table | PRISMA 2020 added items 15 and 22 on the methods and results of certainty assessment | PRISMA 2020; Core GRADE 2025 |
Risk of bias
| Study type | Recommended tool | How it judges | What to know |
|---|---|---|---|
| Randomised controlled trials | RoB 2 (2019) | Five domains: randomisation process, deviations from intended interventions, missing outcome data, outcome measurement, selective reporting; overall low risk, some concerns or high risk | The older Cochrane risk of bias tool (seven items, 2011) is still common in Chinese papers and reviewers may ask for RoB 2; RoB 2 first requires deciding whether the effect of interest is assignment (ITT) or adherence |
| Non-randomised studies of interventions (cohorts or case-control studies comparing treatments) | ROBINS-I (2016); V2 released in 2025 | Uses a target trial as reference and assesses seven domains including confounding, selection and classification of interventions; V2 uses algorithms to suggest judgements from signalling questions and covers immortal time bias | ROBINS-I takes longer than NOS but distinguishes whether confounding was adequately handled |
| Exposure studies (for example smoking and disease) | ROBINS-E (2024) | Similar structure to ROBINS-I, for exposures | The tool is new and Chinese examples are few |
| Observational studies (a common choice in Chinese meta-analyses) | Newcastle-Ottawa Scale (NOS) | Selection, comparability and outcome, up to nine stars | Stang 2010 found that NOS items and scoring lack justification and inter-rater agreement is low; report the judgement for each item rather than only the star total, and do not exclude studies by a '≥7 stars = high quality' rule |
| Diagnostic accuracy studies | QUADAS-2 | Risk of bias and applicability in four domains | Diagnostic meta-analysis uses different pooling methods (bivariate models) from the intervention meta-analysis on this page |
Effect measures
| Data type | Options | How to choose | Basis |
|---|---|---|---|
| Binary | RR, OR, RD | Use RR or OR by default. Relative effects are more consistent across baseline risks than absolute effects, and pooled RDs tend to be more heterogeneous; the Handbook advises against pooling RDs unless there is reason to expect them to be consistent. An OR cannot be read as an RR, and the two diverge noticeably when event rates exceed 10%. Absolute effects in the SoF table are obtained by applying the RR to a baseline risk | Handbook 10.4.3 |
| Continuous, same unit | MD (mean difference) | Use MD for the same scale and unit; results read directly in clinical units. Change scores and final values can be pooled together | Handbook 10.5.2 |
| Continuous, different scales | SMD (Hedges g) | Use SMD only when studies measure the same construct on different scales. The SMD default in RevMan, meta and Stata is Hedges g (with small-sample correction). Change scores and final values cannot be mixed in an SMD because their SDs mean different things. When scales run in opposite directions, multiply one side by -1 first | Handbook 6.5.1.2, 10.5.2 |
| Survival | HR | Pool ln(HR) and its SE (generic inverse variance). Do not replace HR with survival at one time point unless all studies report only that time point. Data sources in order of preference are in the next section | Handbook 6.8.2; Tierney 2007 |
| Counts and rates | Rate ratio | Time units cancel out in a rate ratio, so harmonising them does not change the result | Handbook 6.7.1 |
Data conversion
Ask the corresponding author for the data first. If that fails, convert as below, and remove all converted studies once in a sensitivity analysis to see whether the conclusion changes.
The team behind the Wan, Luo and Shi methods (Tong Tiejun and colleagues, Hong Kong Baptist University) provides an online calculator and Excel sheet that test for skewness before converting; R meta (the median, q1, q3, min and max arguments of metacont) and metafor (conv.fivenum) implement the same formulae.
Handbook 10.5.3 gives a skewness check: for measures that can only be positive, a mean/SD ratio below 2 suggests skew and below 1 indicates clear skew; use the reported mean and SD of such studies in an MD analysis with caution, and do not mix log-transformed and untransformed data in one meta-analysis.
Harmonise units before pooling. For example, divide glucose in mg/dL by 18 to get mmol/L, and total cholesterol in mg/dL by 38.67. Pooling with SMD does not replace unit conversion: SMD treats SD differences caused by different units as differences in effect.
# With only medians and quartiles, enter them directly in metacont; meta converts internally
# Defaults: mean by Luo 2018; SD by Shi 2020 (when the range is available) or Wan 2014 (IQR only)
d <- data.frame(study = c("A", "B", "C"),
n.e = c(40, 60, 35), med.e = c(12, 15, 11), q1.e = c(8, 10, 7), q3.e = c(20, 24, 18),
n.c = c(42, 58, 33), med.c = c(16, 19, 14), q1.c = c(11, 13, 9), q3.c = c(25, 30, 22))
mq <- metacont(n.e, median.e = med.e, q1.e = q1.e, q3.e = q3.e,
n.c = n.c, median.c = med.c, q1.c = q1.c, q3.c = q3.c,
data = d, studlab = study, sm = "MD")
cbind(mq$mean.e, mq$sd.e) # study A: 13.42, 9.21
# For skewed data, switch to the McGrath 2020 Box-Cox method (requires estmeansd)
update(mq, method.mean = "BC-McGrath", method.sd = "BC-McGrath")
# Equivalent metafor function
conv.fivenum(q1 = q1.e, med = med.e, q3 = q3.e, n = n.e, data = d)| What the paper reports | Method | Assumptions and limits |
|---|---|---|
| Median + quartiles (Q1, Q3) | Mean: Luo 2018; SD: Wan 2014 (IQR divided by a factor that depends on sample size and approaches 1.35 for large n) | These methods assume approximate normality. Skewed data lead to underestimated means and SDs; test skewness with Shi 2023 and, if skew is clear, use the McGrath 2020 Box-Cox or quantile estimation methods |
| Median + minimum and maximum | Mean: Luo 2018; SD: Wan 2014 | Handbook 6.5.2.6 recommends not estimating SDs from ranges, which are driven by extreme values |
| Median + quartiles + range (five-number summary) | Mean: Luo 2018; SD: Shi 2020 | Uses the most information and gives the smallest error |
| Mean + 95% CI | SD = √n × (upper − lower) / (2 × t₀.₉₇₅,ₙ₋₁); for n≥60 use 1.96 for t | Valid only when the CI was computed from the t distribution |
| Two group means + between-group P or t value | Back-calculate the between-group SE from t, then the pooled SD of both groups | When a paper only says P<0.05, using P=0.05 overestimates the SD and makes the result conservative |
| Only subgroup n, means and SDs | Combine into one group with the formulae in Handbook table 6.5.a | Do not include the two subgroups as two separate studies |
| Cox regression HR with 95% CI | Use ln(HR) and SE = (ln upper − ln lower)/3.92 directly | First choice |
| Only log-rank P value and number of events | Tierney 2007: variance of ln(HR) ≈ 4/total events (1:1 allocation) | Requires the direction of effect and the allocation ratio |
| Only Kaplan-Meier curves | Digitise with WebPlotDigitizer or Engauge Digitizer, then use the Tierney spreadsheet (Parmar method) or the Guyot 2012 algorithm to reconstruct individual data and fit a Cox model; the R package IPDfromKM implements the reconstruction after digitising | Guyot 2012 validation shows that HRs are estimated reasonably only when the numbers at risk or total events are reported; without both, the reconstructed data assume no censoring, events are overestimated and the CI is too narrow |
Pooling model
Handbook section 10.10.4 is explicit: the choice between fixed and random effects must not be based on a statistical test for heterogeneity. 'Use random effects if I²>50%, otherwise fixed effects' is the most common wrong practice in Chinese tutorials.
| Decision | Recommendation | Basis and conditions |
|---|---|---|
| Fixed or random | Report random effects by default; state this in the protocol and run fixed effects as a sensitivity analysis | Handbook 10.10.4.1 gives no universal recommendation but rejects choosing by heterogeneity test. The two models weight studies differently: random effects give small studies more weight, so the results separate when small-study effects exist |
| tau² estimator | REML | Since the 2024 update the Handbook uses REML as the default (10.10.4.4), and RevMan has defaulted to REML since January 2025. Langan 2019 compared nine estimators and recommended REML; Veroniki 2016 found Paule-Mandel performs well for binary and continuous data. DL underestimates tau² with many small studies or rare events |
| Confidence interval | Hartung-Knapp (HKSJ, t distribution with k−1 df) when tau²>0 and there are more than two studies | In IntHout 2014 simulations the type I error of DL could exceed 30% while HKSJ at most doubled the nominal rate; among 689 Cochrane meta-analyses, 25.1% of results significant with DL were not significant with HKSJ |
| Only two studies | Do not use HK; use a Wald-type interval and compare methods in a sensitivity analysis | With k=2 the t distribution has 1 df and the interval is extremely wide (tested on this page: RR CI 0.004–19.2) |
| tau²=0 or Q < k−1 | Use the Wald-type interval or the modified HK (mKH) | In this situation the HK interval can be narrower than Wald (tested on this page). Röver 2015 recommends the modified version with few studies of varying precision |
| Prediction interval | Report when there are ≥5 studies and no clear funnel plot asymmetry | Handbook 10.10.4.3; use k−1 df (previously k−2). The prediction interval answers where the true effect of a new similar study is likely to fall |
| CI for tau² | Q-profile method | Informative only with ≥5 studies |
Software differences
Most mismatches come from different defaults. In the methods section, state the tau² estimator, the CI method, the distribution and df of the prediction interval, and how I² was computed.
Tested on this page (dat.bcg, 13 studies, REML): the default prediction interval in meta is RR 0.136–1.762 (t distribution, k−1); with Stata's K−2 df it is 0.134–1.785; metafor with test = "z" gives 0.155–1.549. For the same data, I² computed from Q in meta is 92.1% (95% CI 88.3%–94.7%), and computed from tau² in metafor is 92.2% (CI 81.9%–97.7%).
Problems when copying old tutorials: from meta 6.0, hakn = TRUE became method.random.ci = "HK"; comb.fixed became fixed in 5.0 and common in 5.5; byvar became subgroup in 5.0. In 8.5 the old arguments still run but give a deprecated warning; the file argument of forest() was renamed filename in 8.5.
| Setting | RevMan (from January 2025) | Stata 16+ meta | R meta 8.5 | R metafor 5.0 |
|---|---|---|---|---|
| Default tau² estimator | REML (old RevMan 5 used DL) | REML | REML | REML |
| Random-effects CI | Prompts HKSJ when tau²>0 and k>2 | Wald; add se(khartung) for HK | Wald (method.random.ci = "classic"); change to "HK" | z test; test = "knha" for HK |
| Prediction interval | t distribution, k−1 | predinterval option, t distribution, K−2 | prediction = TRUE, t distribution, k−1 (from 8.0; k−2 before) | predict(): normal distribution with test = "z", t distribution with test = "knha" |
| I² formula | From tau² (new) | From tau² in random-effects models | From Q by default; method.I2 = "tau2" to change | From tau² |
| Zero events | MH, Peto; 0.5 continuity correction | mhaenszel, peto; correction value configurable | MH without correction by default; Peto and GLMM available | rma.mh, rma.peto, rma.glmm |
| Egger test / trim-and-fill | Not available | meta bias, egger; meta trimfill | metabias; trimfill | regtest; trimfill |
| Meta-regression | Not available | meta regress | metareg | rma(mods = ~ …) |
Heterogeneity
Subgroup analysis: specify subgroup variables in the protocol; judge by the test for subgroup differences and do not compare P values within subgroups (one significant and one non-significant subgroup does not show a difference). With very few studies per subgroup, within-subgroup random effects can hardly be estimated. In this page's test, splitting by allocation method into three groups gave P<0.0001 for subgroup differences under the common-effect model and P=0.39 under the random-effects model, opposite conclusions; the subgroup with two studies had an HK CI of 0.016–20.8.
Meta-regression: generally not done with fewer than 10 studies, and roughly 10 studies are needed per covariate (Handbook 10.11). The relationship between a study-level covariate (such as mean age) and the outcome may not reflect the individual-level relationship, which is ecological bias. Use the Knapp-Hartung adjustment for meta-regression P values (test = "knha" in metafor, se(khartung) in Stata).
Core GRADE (2025) judges inconsistency by the spread of point estimates, the overlap of confidence intervals and on which side of the decision threshold each study falls, not by I² alone.
| Measure | Question it answers | Common misreading |
|---|---|---|
| Q test | Whether observed differences exceed sampling error | With few studies the test has low power, so P>0.10 does not mean no heterogeneity; with many large studies even small differences become significant |
| I² | What share of the total observed variation comes from true between-study differences | I² is relative and depends on study precision: with tau² unchanged, I² rises towards 100% as study sizes grow (Rücker 2008). I²=50% cannot tell you whether effects range from 40 to 60 or from 10 to 90 (Borenstein 2017). The Handbook bands (0–40%, 30–60%, 50–90%, 75–100%) overlap and are only a rough guide |
| tau², tau | Variance and SD of true effects across studies, on the effect-size scale | tau² estimates are unstable with few studies; report the Q-profile CI |
| Prediction interval | Range of the distribution of true effects | A significant pooled effect with a prediction interval crossing the null means the effect may be absent in some settings; this belongs in the discussion |
Publication bias
Tested on this page (dat.bcg, k=13): Egger P=0.19, Harbord P=0.32, Peters P=0.23, none significant. Trim-and-fill with meta defaults (ma.common = TRUE, iterating with a common-effect model) imputed 4 studies and moved the RR from 0.49 (0.34–0.70) to 0.71 (0.46–1.10), no longer significant; with ma.common = FALSE or metafor's trimfill on an rma object it imputed 1 study, RR 0.52 (0.37–0.74). The heterogeneity in these data comes mainly from latitude (see meta-regression), which is the situation described by Peters 2007.
| Method | When to use | Limitations |
|---|---|---|
| Funnel plot | At least 10 studies; not when studies are of similar size | Asymmetry shows small-study effects, which may come from publication bias, poorer quality of small studies, true heterogeneity, artefactual correlation between effect size and SE, or chance (Handbook table 13.3.b). Contour-enhanced funnel plots show whether the missing region lies in the non-significant zone |
| Egger regression test | MD for continuous outcomes; ≥10 studies | For OR and SMD the effect size is correlated with its SE, so the original Egger test gives false positives; power is low with few studies |
| Harbord / Peters tests | Binary OR | Peters uses 1/total sample size as the predictor, avoiding the artefactual correlation with SE |
| Trim-and-fill | Sensitivity analysis only | Assumes asymmetry is entirely due to publication bias; with substantial heterogeneity it may impute studies and underestimate the effect even without publication bias (Peters 2007). Defaults differ between software and results can differ greatly (see this page's test) |
Sensitivity analysis and GRADE
Certainty ends up as high, moderate, low or very low; build the summary of findings table with GRADEpro GDT.
| GRADE factor | Direction | How to judge |
|---|---|---|
| Starting point | — | RCTs start at high, observational studies at low (when ROBINS-I is used they may start at high and be rated down for bias) |
| Risk of bias | Down | Whether heavily weighted studies are at high risk |
| Inconsistency | Down | Point estimates spread and CIs do not overlap, not explained by prespecified subgroups |
| Indirectness | Down | Population, intervention, comparator or outcome differ from the question |
| Imprecision | Down | The 95% CI crosses the minimal important difference (MID) or the null (Core GRADE 2) |
| Publication bias | Down | Asymmetric funnel plot, or only small industry-funded studies |
| Large effect, dose-response, confounding that would reduce the effect | Up | Mainly for observational studies |
- Change the model: fixed and random effects; tau² by REML, DL and PM.
- Remove studies at high risk of bias (high risk in RoB 2, serious or critical in ROBINS-I).
- Remove studies with converted data (means estimated from medians, HRs reconstructed from curves).
- Handle zero-event studies with another method (see common errors).
- Leave-one-out: remove each study in turn and report the range of pooled effects and whether any single study changes the conclusion; use it to find influential studies, not to pick the subset that lowers I².
- For the primary outcome, analyse adjusted effects only and unadjusted effects only.
Common errors
Double counting a shared control group in multi-arm trials
When doses A and B are both compared with the same control and both comparisons enter one meta-analysis, the control participants are counted twice, their weight is inflated and the two results are correlated. Handbook 23.3.4 recommends combining the relevant experimental groups into one; splitting the control group only partly solves the problem.
Including several reports of the same population
Main reports, subgroup reports and long-term follow-ups often have different author orders and titles. Check registration numbers, recruitment periods and sample sizes, and include each trial once.
Mixing adjusted and unadjusted effects
In observational studies, some report multivariable-adjusted HRs and others only crude HRs or raw counts; pooling them treats results with different confounding control as one effect. Specify in the protocol that the most fully adjusted result is preferred, and run a sensitivity analysis with adjusted results only.
Mishandling zero events
With zero events in one arm, inverse variance methods need a continuity correction, and a 0.5 correction pulls results towards the null. Studies with zero events in both arms are dropped by default for OR and RR. With event rates below 1% and similar group sizes the Peto method has the least bias; with unbalanced groups or large effects use MH without correction or a GLMM (Bradburn 2007). In this page's test on the rosiglitazone data, the choice of method changes the conclusion.
Taking SE for SD
Using 'mean ± SE' from a table as the SD underestimates the SD by a factor of √n and inflates that study's weight and SMD. SD = SE × √n. Check table footnotes during extraction.
Inconsistent units and directions
MDs in different units cannot be pooled directly; if scale directions (higher is better versus higher is worse) are not aligned, the pooled effect cancels out.
Removing studies to lower I²
Deleting studies one by one until I² falls below 50% and reporting the rest is outcome-driven selection. Explain heterogeneity with prespecified subgroups or meta-regression; what cannot be explained goes into the discussion and leads to rating down in GRADE.
Funnel plots and meta-regression with too few studies
An Egger test on 5–6 studies with P>0.05 is reported as 'no publication bias'. With fewer than 10 studies, state that the test was not done and why.
R code
The code runs as is; the example data install with the metadat package (install.packages(c("meta", "metafor"))). For your own binary outcome data you need four columns: events and totals for each group.
# R 4.5.3 + meta 8.5-0 + metafor 5.0.1; example data dat.bcg from the metadat package (13 BCG vaccine trials)
library(meta); library(metafor)
data(dat.bcg, package = "metadat")
bcg <- dat.bcg; bcg$study <- paste(bcg$author, bcg$year)
# 1. Pooling: REML for tau2, both Wald and Hartung-Knapp CIs, plus a prediction interval
m <- metabin(tpos, tpos + tneg, cpos, cpos + cneg, data = bcg, studlab = study,
sm = "RR", method = "Inverse", method.tau = "REML",
method.random.ci = "HK", prediction = TRUE)
summary(m)
update(m, method.random.ci = "classic") # Wald-type interval for comparison
# 2. Sensitivity analysis across tau2 estimators
for (mt in c("DL", "REML", "PM")) print(update(m, method.tau = mt, method.random.ci = "HK"))
# 3. Forest plot (from meta 8.5 the file argument is called filename)
forest(m, sortvar = TE, prediction = TRUE, print.tau2 = TRUE,
filename = "forest_bcg.png")
# 4. Funnel plot and small-study tests (only with k >= 10; use Harbord or Peters for OR outcomes)
funnel(m, type = "contour", common = FALSE)
metabias(m, method.bias = "Egger")
metabias(m, method.bias = "Peters")
# 5. Trim-and-fill: state the model used in the iterations, otherwise meta and metafor disagree
trimfill(m, ma.common = FALSE) # random-effects iterations, same as metafor
trimfill(m) # meta default: ma.common = TRUE (common-effect iterations)
# 6. Subgroups and meta-regression (meta-regression only with >= 10 studies)
update(m, subgroup = alloc)
metareg(m, ~ ablat)
# 7. Leave-one-out
metainf(m, pooled = "random")
# 8. metafor syntax: Hartung-Knapp via test = "knha"; the prediction interval then uses the t distribution
es <- escalc(measure = "RR", ai = tpos, bi = tneg, ci = cpos, di = cneg, data = bcg)
r <- rma(yi, vi, data = es, method = "REML", test = "knha")
predict(r, transf = exp)
regtest(r); trimfill(r); leave1out(r); confint(r)
# If REML fails (Fisher scoring algorithm did not converge):
# rma(yi, vi, data = es, control = list(stepadj = 0.5, maxiter = 1000))Tested on this page
Conditions: R 4.5.3, meta 8.5-0, metafor 5.0.1, metadat 1.6-0, IPDfromKM 0.1.10, macOS arm64. Data are dat.bcg (13 BCG vaccine trials, RR) and dat.nissen2007 (42 rosiglitazone trials, myocardial infarction) shipped with metadat.
| Comparison | Result | Note |
|---|---|---|
| tau²: DL / REML / PM / ML | 0.309 / 0.313 / 0.318 / 0.280 | Pooled RR is about 0.49 in all cases; with 13 studies the estimator has little effect on the point estimate |
| Wald vs HK (REML) | RR 0.489, CI 0.344–0.696 (P=7×10⁻⁵) vs 0.330–0.726 (P=0.002) | HK widens the CI and increases P about 27-fold |
| Weights under fixed and random effects | Largest study TPT Madras: 41.4% fixed, 10.2% random | Fixed-effect RR 0.65, random-effects 0.49; the gap comes from stronger effects in small studies |
| Two studies | Wald 0.134–0.527; HK 0.004–19.2 | Do not use HK with k=2 |
| Four studies, tau²=0, Q=1.04<3 | Wald 0.188–0.310; HK 0.190–0.307; modified HK 0.160–0.363 | HK is narrower than Wald; the modified version corrects this |
| Meta-regression (latitude) | ln(RR) falls 0.029 per degree of latitude; R²=75.6%, residual I²=68% | With KNHA adjustment P goes from <0.0001 to 0.005 |
| Leave-one-out | RR 0.452 (without TPT Madras) to 0.533 (without Hart & Sutherland) | No single study changes the conclusion |
| Zero events (42 trials, 4 double-zero, 26 single-zero) | MH OR 1.28 (0.95–1.73); Peto 1.43 (1.03–1.98); IV + 0.5 correction 1.29 (0.94–1.76); GLMM 1.43 (1.03–1.98) | Peto and GLMM are significant, MH and IV are not; several trials in these data used 2:1 allocation, so the balance assumption of Peto is not met; report several methods side by side |
| SMD: Hedges g vs Cohen d | −0.536 vs −0.544 (9 studies) | Differences arise mainly in small studies with n<30 |
| Median to mean (lognormal, n=200, true mean 23.3, SD 15.8) | Wan/Luo: 20.4, 11.5; Box-Cox (McGrath): 22.5, 13.3 | With skewed data the normal-based methods underestimate the SD by about 27% |
| HR reconstructed from KM curves (simulated data, true HR 0.568) | With risk table: 0.572 (0.417–0.785), events 160/163; without: 0.565 (0.435–0.735), events 231/163 | Without numbers at risk the events are overestimated by 42% and the CI is too narrow, consistent with Guyot 2012 |
Chinese community
The items below come from a Chinese methods journal article, CSDN columns and tools provided by method authors, and only parts that agree with the Cochrane Handbook or the original methods papers are included. The many templated 2025–2026 'meta-analysis pitfall guides' on CSDN were not used.
'Deep extraction' of continuous data
Liu Haining and colleagues (Chinese Journal of Evidence-Based Medicine, 2017, vol. 17, no. 1) cover seven situations: estimating from the median with extremes or quartiles, reading values from figures, combining subgroups, back-calculating the SD from a 95% CI of the mean, and back-calculating a pooled SD from a between-group P or t value, with Stata and Excel commands for t quantiles (invttail, T.INV.2T). They note that using 0.05 when only P<0.05 is reported overestimates the SD. This agrees with Handbook chapter 6.
Extracting HRs with the Tierney spreadsheet
A CSDN column (February 2024) gives the steps after digitising with Engauge Digitizer: enter coordinates and follow-up in Sheet 2a, compare the simulated and original curves in Sheet 2b and re-digitise if they differ, and read the HR and CI in Sheet 4. This agrees with the instructions in the Tierney paper.
Tong Tiejun's online calculator tests skewness first
The online calculator from the authors of Wan 2014, Luo 2018 and Shi 2020 (Zeng Xiantao is also a Shi 2020 author) asks you to click 'Detect the skewness' first, using the Shi 2023 method, and warns against normal-based methods when skew is significant. Many Chinese tutorials quote the formulae and skip this step.
'Use random effects when I²>50%' and 'nine studies are enough for a funnel plot'
Several CSDN and WeChat tutorials repeat these two rules. Handbook 10.10.4.1 explicitly rejects choosing the model by heterogeneity test, and chapter 13 requires at least 10 studies for funnel plot tests. Reviewers often point out both.
Forest plot weights do not depend only on sample size
The Zhihu article ranked first on Bing for 'forest plot' (森林图) says 'the larger the sample, the larger the weight'. This holds only approximately under the fixed-effect model; the random-effects model adds tau² to each study's variance and pulls down the weight of large studies (in this page's test the largest study fell from 41.4% to 10.2%). When reading a forest plot, check which model each weight column belongs to.
Hand to an agent
Search strategy drafting, data conversion, pooling and plotting can be handed to the Scientify scientific agent; inclusion decisions and risk of bias judgements still need you and your co-reviewers.
Example instruction: “Following the attached PROSPERO protocol, run the meta-analysis on the extraction sheet of 12 RCTs: RR for the primary outcome, random effects + REML + Hartung-Knapp, with a prediction interval; convert medians and quartiles with Luo/Wan after a skewness test; run the planned subgroups, leave-one-out and a sensitivity analysis without high risk of bias studies; output forest plots, funnel plots and a draft methods section.”
The agent installs R meta, metafor and other packages on an isolated cloud computer, reads the extraction sheet, checks for SE/SD confusion, inconsistent units and double-counted multi-arm trials, completes conversion, pooling and all sensitivity analyses, and outputs figures and a methods section mapped to PRISMA 2020 items. The agent reviews its results adversarially, for example checking whether REML and DL, HK and Wald agree, and whether the trim-and-fill iteration model is stated. The workspace keeps scripts, parameters, logs and data for reproduction, and the task keeps running after you close your computer.
You still need to check: inclusion and exclusion decisions, risk of bias judgements, whether the direction and extracted values of effects match the papers, and the reasons for GRADE ratings.
Sources
- Cochrane Handbook 6.5 chapter 10: Analysing data and undertaking meta-analyses — 10.4.3 effect measures; 10.4.4 zero events; 10.5.2 change scores and SMD; 10.10.2 I² bands; 10.10.4 REML, HKSJ, prediction intervals; 10.11 subgroups and meta-regression
- Cochrane Handbook 6.5 chapter 6: Choosing effect measures and computing estimates — Hedges g; IQR and SD; no SDs from ranges; extracting HRs (6.8.2)
- Cochrane Handbook 6.5 chapter 13: Bias due to missing results — At least 10 studies for funnel plot tests; other causes of asymmetry; Egger not for OR and SMD
- Cochrane Handbook 6.5 chapter 23: Special designs including multi-arm trials — Double counting of shared control groups and how to handle it (23.3.4)
- Cochrane Statistical Methods Group newsletter, December 2025 — RevMan update of 23 January 2025: HKSJ, REML default, Q-profile, new I² formula, prediction interval with k−1
- PRISMA 2020 statement (Page et al., BMJ 2021) — 27 items; item 7 full search strategies; item 8 number and independence of screeners; new certainty items
- PROSPERO new website announcement (March 2025) — Automatic processing and publication, search for similar topics
- IntHout et al. 2014: HKSJ versus DL — DL error rates can exceed 30%; 25.1% of significant results not significant with HKSJ
- Langan et al. 2019: simulation comparison of nine tau² estimators — REML recommended; DL negatively biased with small studies and rare events; PM positively biased with very different study sizes
- Veroniki et al. 2016: review of between-study variance estimators — PM (binary and continuous) and REML (continuous) recommended; Q-profile for the tau² CI
- Röver et al. 2015: HKSJ and its modification — Modified HK recommended with few studies of varying precision
- Rücker et al. 2008: undue reliance on I² may mislead — I² rises with study precision at constant tau²
- Borenstein et al. 2017: I² is not an absolute measure of heterogeneity — Use the prediction interval to express the range of effects
- Peters et al. 2007: trim-and-fill under heterogeneity — May impute studies wrongly with large heterogeneity; use only as sensitivity analysis
- Bradburn et al. 2007: methods for rare events — Peto best with event rates <1% and balanced groups; otherwise MH without correction or logistic regression
- Wan et al. 2014, Luo et al. 2018, Shi et al. 2020: estimating mean and SD from medians — Luo 2018 in Stat Methods Med Res 27:1785; Shi 2020 in Res Synth Methods 11:641
- McGrath et al. 2020: estimating mean and SD for non-normal data — Box-Cox and quantile estimation methods; R package estmeansd
- Tong Tiejun group online calculator (median2mean) — Shi 2023 skewness test first, then Luo/Wan/Shi conversion
- Tierney et al. 2007: incorporating summary time-to-event data into meta-analysis — Estimating HRs from P values, event counts and survival curves, with Excel spreadsheet
- Guyot et al. 2012: reconstructing data from published KM curves — HRs require numbers at risk or total events to be accurate
- IPDfromKM (Liu et al. 2021) — R package and Shiny app: digitising, reconstruction and accuracy checks
- RoB 2 (Sterne et al. 2019) — Five risk of bias domains for RCTs
- ROBINS-I V2 introduction (Cochrane, 2025) — Algorithmic judgements, immortal time bias
- Stang 2010: critical evaluation of the Newcastle-Ottawa Scale — Arbitrariness of NOS scoring
- Core GRADE 2 and 3 (BMJ 2025) — Imprecision judged against the MID; inconsistency by point estimates and CI overlap
- R package meta NEWS — Prediction interval k−1 from 8.0, new method.I2; forest argument file renamed filename in 8.5; hakn replaced by method.random.ci in 6.0
- metafor: convergence problems with rma — Fisher scoring algorithm did not converge; stepadj, maxiter, switch to PM
- Stata meta summarize manual — REML default; se(khartung); prediction interval with t distribution, K−2
- Experience post: deep extraction when papers lack means and SDs (CSDN, summarising Liu et al. 2017) — SD from CI, P values and subgroups; valid for normal data or CIs from the t distribution
- Experience post: extracting HRs from survival curves with Engauge Digitizer and an Excel template (CSDN) — How to use each sheet of the Tierney spreadsheet