Use this before submission when the evaluation is not yet locked. SoCC reviewers come from both SIGMOD and SIGOPS, and the evaluation is where a good cloud idea is won or lost. The organizing principle is measured evidence proportional to the claim — the study must test the thing the paper asserts, on a system and workloads a skeptic from either community would accept, with tail latency and cost reported, not just the mean.
| Cloud claim | Matching evidence | Reject pattern avoided |
|---|---|---|
| "System raises throughput at scale" | Throughput across node counts on a real testbed vs. tuned baseline | "Simulated / small-scale only" |
| "Holds tail latency under bursts" | p99 distribution under bursty replayed load | "Only mean latency reported" |
| "Cuts cost" | Instance-seconds or $ under a stated pricing model | "Cost asserted, never measured" |
| "Fair under multi-tenancy" | Per-tenant SLO attainment with contending workloads | "Single-tenant microbenchmark" |
| "The new mechanism adds the value" | Ablation isolating the mechanism vs. a simpler policy | "Mechanism's marginal value never isolated" |
[Tail] report percentiles and the distribution, not just the mean; state the SLO target
[Cost] define the pricing model; report the cost metric so a reader can recompute the saving
[Scale] sweep nodes/tenants/load; identify the bottleneck where the result degrades
[Runs] report the number of runs and the variance for any measured metric
[Compute] state the testbed actually used (nodes, instance types, kernel), not vague feasibility
Suppose the paper claims a scheduler enforces latency SLOs on shared storage better than the prior system. The matching plan: run both on a real storage testbed with a mix of latency-sensitive and batch tenants replayed from a representative workload; tune both with an equal, documented budget; report per-tenant p99 attainment and aggregate throughput with variance; show behavior as tenants and load scale and name the bottleneck; report the provisioning cost of each; and state external-validity limits (storage backend, workload mix) — every number traceable to a logged run in the artifact.
[Evaluation readiness] strong / adequate / weak
[Claim -> evidence map] <claim: testbed/workload/metric incl. tail+cost>
[Deployment realism] <real testbed / scaled deployment / simulation-only>
[Baseline fairness] <baseline -> tuned? equal budget? documented?>
[Tail+cost+scale] <percentiles + cost model + scaling behavior reported? yes/no>
[Provenance] <trace SHAs+dates, replay harness, testbed described? yes/no>
[Decision-critical next run] <one experiment or measurement extension>