Use this before submission when the empirical story is not yet locked. FAccT evidence is not leaderboard evidence: the reviewer pool asks whether your study actually shows the harm, disparity, accountability gap, or transparency effect you claim, on the people you claim it for, with methods honest about their limits. The organizing principle is evidence proportional to the claim — and because FAccT is interdisciplinary, "evidence" can be a disaggregated statistical audit, a coded interview corpus, a participatory study, or a documented case, each held to its own field's standard of rigor.
| FAccT claim | Matching evidence | Reject pattern avoided |
|---|---|---|
| "System X harms group G more" | Disaggregated error/outcome metrics by G, with CIs, on real data | "Aggregate accuracy hides the subgroup gap" |
| "Our method reduces the disparity" | Gap before/after vs. a tuned fairness baseline + the accuracy cost | "Fairness improved, utility cost never reported" |
| "Affected people cannot contest decisions" | Interviews/observation with those people, coded and reflexive | "Researcher speculation stands in for lived experience" |
| "This documentation improves transparency" | A study of whether real users act differently with it | "Assumed usefulness; never tested with a reader" |
| "The proxy is valid for the protected attribute" | Validation of the proxy against ground truth, error stated | "Proxy treated as truth; construct threat ignored" |
[Definition] state how each group is defined; whose categories are these, and who is erased by them?
[Measurement] is the attribute observed, self-reported, or inferred? report proxy error and bias
[Intersection] test intersectional subgroups where numbers allow; note where they are too small
[Consent] do the people classified know and agree? document the ethics basis
[Missingness] who is absent from the data entirely, and how does that bound the claim?
A paper claims a screening model disadvantages a protected group. The matching plan: obtain or construct a realistic labeled dataset with documented provenance; report selection/error rates disaggregated by group and intersection with confidence intervals; validate the group proxy and state its error; compare against the vendor's fairness claim; audit a sample of individual cases qualitatively for face validity; document the consent/ethics basis for using the data; and state plainly which populations the audit cannot speak to — every number traceable to a logged analysis in the supplementary material.
[Evaluation readiness] strong / adequate / weak
[Claim -> evidence map] <claim: population / metric-or-method / uncertainty>
[Disaggregation] <groups reported? proxy validity stated? intersections where possible?>
[Ethics basis] <consent / IRB / compensation / re-harm avoidance documented? yes/no>
[Qual rigor] <coding / agreement / reflexivity present where relevant? yes/no>
[Decision-critical next run] <one study extension or analysis>