Interspeech added an author rebuttal to its historically response-free pipeline in recent cycles; for 2026 the response window was reported as 24 April – 1 May, ahead of the 5 June notification. Confirm the live mechanics (length cap, CMT form fields, whether replies go to reviewers or only the area chair) before writing a word — rebuttal formats at this venue are young and still moving.
An Interspeech review set typically mixes reviewer species, because ISCA's areas span science and engineering:
| Reviewer instinct | What they attacked | What convinces them |
|---|---|---|
| Speech engineer | Baselines stale, WER delta within noise | Point to the significance test or per-condition breakdown already in the paper |
| Speech scientist | No phonetic/error analysis, claims exceed perception data | Cite the analysis subsection; concede scope honestly |
| ML generalist | "Incremental over the arXiv version of X" | State the technical delta in one sentence; note what the 4-page record shows |
| Resource skeptic | Corpus license, split leakage, evaluation protocol | Quote the exact split/protocol line; name the license |
Answer each reviewer in their own dialect. A MOS-methodology objection answered with more WER numbers reads as a dodge.
Unlike venues with appendices, everything reviewable was on those four pages — so "see appendix" is never available. Your levers are:
Whether new numbers are allowed in the response varies by cycle; if unverified, assume tables of new results are unwelcome and offer them as camera-ready additions.
The response is part of the double-anonymous record and the anonymity period runs until decisions. No identity hints, no links to demo pages created after submission under identifiable accounts, no "as our group showed at Interspeech 2024."
One block per objection, ordered by decision weight:
R1-Q2 (statistical validity of the 0.4 WER gain):
- The concern: gains may be within run-to-run variance.
- In the record: Table 2 reports mean±sd over 3 seeds; Sec 4.2 gives the
bootstrap 95% CI on test-other, excluding zero.
- Clarification: the CI uses utterance-level resampling (1000 draws).
- Offered for camera-ready: add the matched-pairs test on test-clean.
Four to six lines each. The meta-reviewer will compress you anyway — pre-compress.
Interspeech reviews arrive compressed and sometimes blunt — three reviewers covering a 4-page paper leave little space for pleasantries. The productive reading: every blunt sentence is a compression artifact of a real concern. Extract the concern, drop the tone, and never mirror bluntness back; the meta-reviewer reads your temperament as evidence about the work's solidity.
Rebuttal mechanics at Interspeech are recent and not guaranteed to recur. If the
current cycle runs without one, this skill still applies at two later moments:
the camera-ready revision memo (accepted papers), where reviewer concerns get
addressed in print, and the next-cycle revision plan (rejected papers), where
the same objection map drives the rewrite — see interspeech-review-process
for the triage lanes.
[Review set] R1/R2/R3(+) one-line objection map, tagged by type
[Decisive objection] <the one the meta-reviewer will weigh>
[Anchors] <section/table/figure of the 4-page record answering each point>
[Concessions] <honest gaps + camera-ready offers>
[Rebuttal draft] <numbered blocks per skeleton, within the live cap>
[Verify first] cap / new-results policy / who reads it — from the live CMT form
The 2026 window dates above came from secondary renderings and are flagged 待核实 in
resources/official-source-map.md; the official author pages control.