agent-evaluation-reporting
sickn33/agentic-awesome-skills
Generate decision-ready reports from AI agent evaluation runs. This skill ensures autonomous, assisted, failed, and timed-out outcomes remain distinct and comparable. It defines metrics with correct denominators, handles latency and cost populations honestly, and maps evidence to decision gates. Ideal for benchmarking, regression testing, and production readiness validation of AI agents.