Use this before submission when the empirical story is not yet locked. ICASSP reviewers are subfield experts who know the right metric and the right baseline for your task, so the fastest route to rejection is the wrong ruler or a stale comparison. The four pages force a small number of decisive experiments, not a large number of weak ones.
| Task | Standard metric(s) | Standard evaluation anchor |
|---|---|---|
| Speech recognition | WER / CER | LibriSpeech, WSJ, or task corpus with fixed split |
| Enhancement / separation | SI-SDR, PESQ, STOI | Matched mixture set, reference-aligned scorer |
| Speaker / language ID | EER, minDCF | Standard trial lists (e.g., VoxCeleb-style) |
| Sound event / audio tagging | mAP, F1, error rate | Fixed labeled set, defined operating point |
| Image / video restoration | PSNR, SSIM | Standard test set, defined borders and depth |
| Communications | BER / BLER vs SNR | Defined channel model and decoder |
| Estimation / detection | RMSE, ROC/AUC | Monte-Carlo trials, bound (Cramér-Rao) if apt |
Reporting the wrong metric family (e.g., classification accuracy for a separation paper) is a first-round reject pattern; match the ruler to the task before anything else.
Fig. 2: metric vs condition (e.g., SI-SDR vs input SNR, 0-20 dB)
- proposed (mean ± sd over 3 seeds)
- strong baseline (same corpus, same scorer)
Table 1: ablation — remove one component at a time, same protocol
- full method | -component A | -component B | baseline
Report: corpus + split, scorer config, seeds, run count, hardware/runtime
A submission claims improved dereverberation. The matching plan: evaluate on a standard reverberant set with PESQ and STOI using a fixed scorer, sweep reverberation time (RT60) rather than reporting one room, ablate the key module, draw a current strong baseline on the same axes, and report the mean and spread over seeds — every panel tied to the claim it supports.
[Experiment readiness] strong / adequate / weak
[Metric fit] task-matched? <metric -> task>
[Baseline] current-strong / standard-corpus? yes/no
[Condition sweep] present over <axis>? yes/no
[Missing evidence] <ablation / spread / baseline / condition>
[Decision-critical next run] <one experiment>