Use this to make an EACL paper's evidence hold up under NLP review. EACL rewards well-scoped
questions answered with careful controls over leaderboard maximalism — its best papers include
analyses and critiques, not only state-of-the-art systems (see
../../resources/exemplars/library.md). Design the evidence to match the claim exactly, no
broader.
| Claim | Required breadth |
|---|---|
| "Works for language L" | Solid results on L, honestly scoped |
| "Cross-lingual / multilingual" | Enough languages across resource levels; per-language results |
| "General method" | Multiple tasks/datasets, not one convenient benchmark |
| "Robust" | Stress tests / shifts, not just in-distribution |
A multilingual claim backed by two high-resource languages is the classic EACL over-reach — the morphology-across-57-languages exemplar shows the bar.
Evidence floor for a headline comparison:
seeds: >= 3-5 runs
report: mean +/- CI (or std), never a lone run
significance: a test when systems are close
ablations: isolate each component's contribution
eacl-artifact-evaluation).[ ] Baselines tuned, search spaces stated
[ ] LLM baselines with verbatim prompts + decoding
[ ] Breadth matches the claim (per-language results if multilingual)
[ ] >= 3-5 seeds; variance/CIs reported
[ ] Significance test where systems are close
[ ] Ablations isolate each component
[ ] Contamination addressed
[ ] Human eval: annotators, agreement, pay reported
[ ] Error analysis quantified
[Evidence strength] Strong / Adequate / Weak
[Baseline fairness] <tuned? LLM baseline reported?>
[Breadth vs claim] <matched / over-reaching>
[Variance + significance] <seeds, CIs, tests>
[Contamination + human eval] <controls present?>
[Fix order] <experiments to add/scope before the cycle deadline>