An ARE review runs no identification of its own — there is no design to defend, no replication package of your data. Instead you act as the field's referee-of-record for adjacent readers: you judge how much weight each primary study can bear so the review weighs the evidence correctly. Make the appraisal explicit, not implicit:
You are not re-running these — you are rating their credibility in the evidence matrix so the review's conclusions track the best evidence, not the loudest paper.
A few moments in an ARE review justify running something instead of only rating it: a survey table pools estimates and wants a formal meta-analytic average or publication-bias check; a pivotal magnitude predates the staggered-DiD corrections and its replication data are public, so you can report what callaway_santanna or bacon_decomposition actually does to it rather than speculate; or a weak-instrument worry could be settled by an effective_f_test on the archived first stage. For these, hand off to execution-with-mcp, which maps each design and reviewer objection to the callable StatsPAI / Stata MCP chain (detect_design → fit with as_handle=true → audit_result). Label any such figure as your re-analysis, and never report a number you did not compute.
Conflicting results are reconciled by credibility and by what each study estimates, never by tallying "7 studies positive, 4 negative." Two estimates that disagree often measure different objects (different populations, estimands, time horizons); say so, and let the framework's cells carry the distinction. A pooled "consensus" across non-comparable designs manufactures false agreement that ARE's methodologically literate readers will catch.
A review must be comprehensive in coverage yet selective in emphasis — and stay accessible in ~25–40 pages. Tier the corpus:
| Tier | Treatment |
|---|---|
| Foundational / field-defining | discussed in text, with what they established and their limits |
| Important contributions | grouped and weighed within framework cells; cited with their finding |
| Confirmatory / incremental | cited in clusters ("see also …") to show coverage without bloating prose |
| Tangential | cited only where they bear on a specific claim |
Comprehensiveness is proven by the citation set + saturation log (arecon-literature-synthesis); selectivity is exercised in the prose.
ARE referees are frequently the surveyed authors themselves, so balance is strategic as well as ethical:
【Credibility appraisal】pivotal studies rated (design/identification/robustness)? Y/N
【Conflict handling】reconciled by credibility + estimand (not vote-count)? Y/N
【Tiering】corpus split foundational/important/confirmatory/tangential? Y/N
【Comprehensiveness】saturation log supports "nothing important missing"? Y/N
【Steelman】each rival school stated at its strongest? Y/N
【Controversy】evidence-to-settle stated; author's read labelled? Y/N
【Self-citation audit】own work at warranted tier; emphasis identity-blind? Y/N
【Next step】→ arecon-tables-figures (who-found-what tables) → arecon-writing-style