"Not convincing" is the standard Interspeech rejection, and it almost always means the experimental design — not the idea — failed. Speech evaluation has decades of conventions per task; an experiment section that ignores them is illegible to the reviewer pool regardless of how good the numbers are.
| Task family | Primary metrics | Convention notes |
|---|---|---|
| ASR | WER / CER | normalization + scorer disclosed; CER for unsegmented scripts |
| TTS / VC | MOS, CMOS (+ objective proxies) | panel protocol reported; CMOS for close systems |
| Speaker verification | EER, minDCF | official trial lists; DCF prior/costs stated |
| Diarization | DER / JER | collar and overlap handling stated |
| Enhancement / separation | PESQ, STOI/ESTOI, SI-SDR (+ DNSMOS-style proxies) | wideband vs narrowband named |
| SLU / speech translation | intent acc / F1, BLEU/COMET on ASR output | cascaded vs end-to-end made explicit |
| Paralinguistics / health | UAR, F1 | speaker-disjoint splits are mandatory |
Using a proxy where the community expects the primary (e.g., only neural MOS predictors for a TTS claim) needs an explicit defense sentence.
State which variation your statistics cover — the two are routinely conflated:
interspeech-reproducibility).A speech claim is implicitly quantified over speakers, acoustic conditions, and often languages. Reviewers probe the quantifier:
interspeech-artifact-evaluation); leaked or scraped audio can sink an
otherwise strong paper on ethics review.Budget roughly one column for the decisive comparison, half for ablation, half for analysis. The analysis half is what separates accepted Interspeech papers: one error-pattern finding (where the gains live — short utterances, overlapping speech, a phone class) converts a benchmark delta into a scientific statement.
Claim: proposed 4.9% vs baseline 5.6% WER on test-other (2939 utts).
1. Paired per-utterance errors → paired bootstrap, 1000 resamples.
2. Δ WER 95% CI: [-0.9, -0.5] — excludes 0 → test-set variation covered.
3. Across 3 seeds: 4.9, 5.0, 4.8 (sd 0.1) vs 5.6, 5.7, 5.6 (sd 0.06)
→ training variation does not swallow the gap.
4. Report: "−0.7 abs. WER (95% CI [−0.9, −0.5], paired bootstrap;
consistent across 3 seeds)."
Two randomness sources, two checks, one sentence in the paper. If step 2's CI had straddled zero, the honest paper reports the trend and softens the verb — and usually survives review better than the inflated version.
Interspeech's mixed jury respects a disclosed regression far more than a suspicious clean sweep. If the method loses on clean speech while winning on noisy, print both numbers and make the trade-off the story — condition-dependent behavior is a finding in a field about acoustic variability, and hiding it is the reviewer-trust equivalent of a failed significance test.
[ ] Primary metric matches task convention; ruler disclosed
[ ] Baseline set includes a public-recipe or challenge anchor
[ ] Each A>B claim carries a CI or matched-pairs test
[ ] Seeds: n stated; variance reported or single-run admitted
[ ] Speaker-disjoint splits verified where required
[ ] Condition/language breakdown present; regressions named
[ ] Dev-only tuning stated; test touched once
[ ] One analysis finding, not just deltas
[Claim inventory] each claim → metric → evidence status
[Metric-law check] conventions met / violations
[Baseline verdict] anchored / self-referential / stale
[Statistics] randomness covered (test-set / seeds / raters) per claim
[Coverage gaps] speaker / condition / language / leakage
[Cheapest decisive fix] <one experiment that most raises conviction>
Metric conventions are community law rather than CFP text and move slowly, but
challenge editions and recipe baselines roll every year — re-anchor at design
time (sources logged in resources/official-source-map.md, 2026-07-08).