Use this before writing. ICASSP is the IEEE Signal Processing Society flagship, and its defining feature is breadth: speech and audio, image and video, communications and radar, sensor arrays, estimation and detection theory, and machine learning for signals all review under one roof. Speech is one track among many, not the identity of the venue. A project fits when its contribution changes a signal-processing primitive and can be proved in four pages.
Name what the contribution actually is:
Speech papers can go to either venue, and picking wrong costs a cycle. They differ on mechanics as much as on scope:
| Dimension | ICASSP (IEEE SPS) | Interspeech (ISCA) |
|---|---|---|
| Scope | All signal processing; speech is one track | Speech and spoken language only |
| Culture | Signal-processing methods, math-forward | Speech-science + engineering, spoken-language focus |
| Format | 4+1 (4 content pages + reference page) | 4+1 (4 content + reference page), different template |
| Anonymity | Single-blind — author list included | Double-anonymous with a pre-deadline anonymity period |
| Deadline | September (autumn) | Late February (winter) |
| Portal | CMS (cmsworkshops.com) |
Microsoft CMT |
| Proceedings | IEEE Xplore | ISCA Archive, open access |
Decision rule: route to ICASSP when the contribution is a signal-processing method whose speech application is one instance (a new front-end, estimator, separation objective, or array technique), or when the September calendar and single-blind culture fit. Route to Interspeech when the contribution is fundamentally about spoken language — phonetics, prosody, dialogue, speech science, or a speech-specific model where the linguistic content is the point — or when the February calendar and double-blind norms fit. When both fit, let deadline and reviewer community decide.
| Signal in the project | ICASSP reading |
|---|---|
| New estimator/filter/detector with a signal model + evaluation | Core fit — the house genre |
| Speech/audio method framed as a signal-processing mechanism | Core fit (vs Interspeech if it is spoken-language-centric) |
| Image/video restoration, coding, or analysis | Core fit (compare ICIP for image-specific work) |
| Communications, radar, array, or sensor signal processing | Core fit |
| Generic deep net with strong benchmarks but no signal insight | Better at NeurIPS/ICML/ICLR/AAAI |
| A finished, journal-length result with extensive theory | An SPS journal (TSP/TASLP/SPL) |
[Fit] strong ICASSP / possible ICASSP / better elsewhere
[Primitive] estimator / transform / representation / objective / resource / model-based-DL / none
[If speech] ICASSP vs Interspeech -> <decision + reason (scope/calendar/blinding)>
[Best venue] ICASSP / Interspeech / ICIP / EUSIPCO / WASPAA / SPS journal / ML venue
[Contribution sentence] <one signal-processing claim>
[Next action] <framing, experiment, or venue switch>