The response you write inside an ARR cycle is read twice: once by the reviewers and action editor who finish the meta-review, and once more — frozen — by NAACL's senior area chairs when you commit the package. That second audience changes the calculus. A reply that merely wins an argument helps less than a reply that leaves a clean, quotable record showing the objection was answered.
ARR response windows are short and reviewers rarely return for a second round. Rank objections by what they cost you downstream:
| Objection class | Downstream cost if unanswered | Response priority |
|---|---|---|
| Soundness attack (flawed eval, wrong baseline, leakage suspicion) | Sinks both ARR scores and any NAACL commitment | Answer first, with evidence coordinates |
| Coverage attack ("claims exceed the languages/domains tested") | Recurring killer for Americas-flavored papers | Concede scope or cite the table that grants it |
| Novelty doubt | Damaging but survivable if excitement holds | Answer with a two-sentence delta statement |
| Presentation complaints | Cheap to fix, cheap to promise | Batch into one short commitment list |
| Requests for infeasible new experiments | Low — ARR forbids judging you on promises alone | Scope honestly; run only what fits the window |
R2 raises leakage: the eval set may overlap pretraining data.
1. What we already did: §4.2 reports n-gram-overlap screening
against the released corpus list (Table 5, appendix C).
2. What we ran this window: exact-match dedup on the 2 newest
models; contaminated items = 0.8%, results shift < 0.3 F1.
3. What we will state at camera-ready: the screening protocol
moves from appendix to §4, with the dedup numbers added.
Three moves — existing evidence, in-window evidence, camera-ready promise — each labeled, none blurred into the others.
The frozen record amplifies tone: a defensive sentence reads worse at commitment than it did in the heat of the window.
| Instinctive draft | What the record shows | Repaired version |
|---|---|---|
| "The reviewer clearly did not read Section 4." | Author hostility, unresolved thread | "Section 4.2 addresses this directly; we suspect the two-column layout buried it, and will promote the key sentence." |
| "This is standard practice in the field." | Appeal to authority, no evidence | "Three recent papers on this task use the same protocol [refs in §2]; we follow them for comparability." |
| "We will add these experiments to the final version." | An unverifiable promise doing a result's job | "We ran the two feasible configurations this week: X and Y (numbers below); the full grid is future work." |
| "We respectfully disagree." (paragraph follows) | Disagreement without a decision path | "The disagreement reduces to one empirical question: Q. Table 3 answers it as follows." |
Before sending, reread each reply pretending you are a NAACL SAC months later, skimming to decide Main versus Findings versus reject. Ask:
A response that fails this test may still sway one reviewer; it will do nothing for the audience that actually admits papers to NAACL.
Because NAACL sorts committed papers into Main and Findings, the response has two distinct win conditions, and knowing which one you are playing for changes the drafting:
State the tier goal in the team's triage notes before anyone drafts; mixed responses that half-defend correctness and half-oversell impact do neither job well.
[Objection map] <reviewer -> class -> priority>
[In-window evidence] <what was actually run, with results>
[Concessions] <exact sentences offered>
[Camera-ready promises] <bounded list>
[Commitment-record check] pass / weak / fail per thread