gaia-debugging
ruvnet/ruflo
This skill provides a systematic workflow for diagnosing failures in GAIA benchmark runs. Users can analyze task traces, classify failure modes (e.g., Tool Gap, Reasoning Miss, Extraction Bug), identify root causes, and apply targeted fixes to improve AI agent performance. It is essential for optimizing LLM benchmarks and finding performance regressions.