Close the last mile: turn "you should use a heterogeneity-robust DiD / weak-IV-robust CI /
multiple-testing correction" into an actual fitted, audited estimate. Full map +
orchestration spine + validated worked-examples (DiD / IV / RDD / synthetic-control / DML):
shared-resources/empirical-methods/execution-with-mcp.md.
detect_design → preflight / recommend → fit with as_handle=true.audit_result(result_id) — enumerate the checks the design still owes; run each
suggest_function it names.honest_did_from_result,
sensitivity_from_result, evalue_from_result).bibtex(keys=[…]) for citations — never invent references.callaway_santanna / sun_abraham + bacon_decomposition + honest_did_from_result.iv + effective_f_test + anderson_rubin_ci (weak-IV-robust).rdrobust + rddensity / mccrary_test.synth / sdid + placebo inference.dml + dml_diagnostics (overlap) + oster / sensemakr.romano_wolf, wild_cluster_bootstrap, twoway_cluster.etable / did_summary_to_latex straight from the handle.bibtex is the only citation source.
official-source-map.md.code/ skeleton
and flag any unverified number.【Design】DiD / IV / RDD / SCM / DML / …
【Estimate】point [CI] (estimator)
【Key diagnostic】(first-stage/effective F, pre-trends p, overlap, placebo p, …)
【Audit gaps run】…
【Magnitude】interpretable units
【Next】rt-submission-readiness / the pack's tables-figures skill