Use this before submission when the evaluation is not yet locked. IPSN reviewers are sensor-systems and information-processing specialists; the evaluation is where a good idea is won or lost. The organizing principle is evidence measured on real hardware against real ground truth — the evaluation must test the sensing claim the paper actually makes, on platforms and baselines a skeptic would accept.
| Sensing claim | Matching evidence | Reject pattern avoided |
|---|---|---|
| "Estimator is more accurate" | Error vs ground truth on real traces, with CIs, vs a tuned baseline / a bound | "Simulated inputs only" |
| "Runs within an energy budget" | Measured µJ/op on an instrumented rail on the real MCU | "Energy estimated from datasheet" |
| "Localizes to X meters" | Surveyed ground-truth positions; error distribution, not just mean | "Ground truth from the same model being tested" |
| "Deploys reliably" | Yield, sync error, packet loss over a real deployment duration | "Idealized single-run numbers" |
| "The on-device model adds value" | Ablation vs classical DSP / heuristic on the same hardware | "Model's marginal contribution never isolated" |
| "Scales to N nodes" | Real or emulated multi-hop at realistic scale, with the bottleneck named | "Two-node bench test, universal claim" |
[Platform] exact MCU/SoC, clock, RAM/flash; the sensor and sampling regime
[Energy] µJ per inference/op, instrument named; duty cycle if always-on
[Latency] end-to-end on-device latency; number of runs and variance
[Footprint] model/pipeline RAM+flash vs available; what had to be quantized/pruned
[Contamination] for learned components, keep train/field data disjoint; report the split
[Ablation] learned component vs DSP/heuristic baseline on the same node
ipsn-reproducibility).The paper claims a new estimator localizes better than the prior method. The matching plan: collect real RF/acoustic traces at surveyed positions; run both estimators on the same traces under equal tuning; report the full error distribution (not just the mean) with confidence intervals; compare against the relevant estimation-theoretic bound; and state the environments (indoor/outdoor, multipath regimes) as a bounded external-validity limit — every number traceable to a logged run and the surveyed ground truth in the artifact.
[Evaluation readiness] strong / adequate / weak
[Claim -> evidence map] <claim: platform / ground truth / metric / statistic>
[Real-hardware check] measured on real sensors/MCU, not simulation only? yes/no
[Energy accounting] <µJ/op measured? instrument named? footprint reported?>
[Baseline fairness] <strongest prior + simple baseline, equal conditions, same hardware?>
[Limits-by-design] <site / calibration / generalization -> instrumentation to bound it>
[Decision-critical next run] <one experiment or deployment extension>