OOPSLA's evaluation question is not "is there a big table?" but "does the
evidence type match the claim type?" The venue's published scope runs
from mathematical formalisms to empirical studies, and its exemplars span
benchmark suites, measurement methodology, corpus mining, and language
experience reports (resources/exemplars/library.md) — so the first design
act is choosing the right instrument, and the second is executing it to the
SIGPLAN Empirical Evaluation Guidelines standard that reviewers apply
checklist-in-hand (oopsla-reproducibility operationalizes the pillars).
| Claim type | Primary evidence | Common OOPSLA failure |
|---|---|---|
| "Faster / cheaper" | Benchmarks vs strongest baseline, variance reported | Weak baseline; startup vs steady-state conflated |
| "More expressive / safer" | Formal result + programs witnessing the boundary | Expressiveness asserted by example only |
| "Programmers benefit" | User study or field data with a design | Anecdote from the authors' own use |
| "Occurs in practice" | Corpus study with stated selection rule | Convenience sample of famous repos |
| "The design generalizes" | Second instantiation (language/runtime/domain) | Single-host generalization claims |
| "Semantics is right" | Mechanization or proofs + conformance tests | Calculus untethered from the implementation |
A paper may need two rows; it rarely supports five. Cutting a claim is cheaper than defending its missing evidence through a Major Revision.
The two-round system changes experimental economics. Between an R1 verdict and the R2 resubmission there are only months, so:
Design now, before Round N:
- matrix of runs a reviewer could plausibly demand (extra baseline,
larger corpus, second platform) with wall-clock + hardware cost each
- keep the harness parameterized so a demanded cell is a config change
- archive raw results per run (the ledger of oopsla-reproducibility)
Payoff: a Minor Revision executes in days; a Major Revision's
expectation list maps to known cells instead of new engineering.
oopsla-writing-style).[Routing] claim → evidence rows used + mismatches found
[Baseline audit] strongest-sensible test: pass / gaps
[Workload rule] stated / absent; exclusions justified: yes/no
[Demand matrix] anticipated reviewer demands with cost estimates
[Analysis floor] repetitions/dispersion/summary-statistic compliance