Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Diagnose why a GAIA question failed — extract trace, classify failure mode, and propose a fix. Use when a GAIA benchmark run reports a failed/incorrect task_id and you need to root-cause it before resubmitting.
.claude/skills/ruvnet-gaia-debugging/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-08 | ✗→✓ | ▲ Improved | 22% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 13% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -1% | 0% |
| case-01 | ✗→✓ | ▲ Improved | 773% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 960% | 0% |
When a GAIA question fails, systematically diagnose the root cause and propose a targeted fix.
task_id returns the wrong answer or times out| Code | Mode | Symptom | Fix direction | |------|------|---------|--------------| | TG | Tool Gap | Agent lacks a required tool (no image OCR, no PDF reader) | Add tool to catalogue | | RM | Reasoning Miss | Agent has the right data but draws wrong conclusion | Improve system prompt, add CoT instruction | | EB | Extraction Bug | Answer is in the trace but FINAL_ANSWER: regex fails | Fix answer extraction pattern | | LI | Loop Issue | Agent loops (re-asks same tool call) and hits turn limit | Increase max-turns or add loop-detection | | DS | Dataset Shift | Ground truth differs from what web currently shows | Flag for HAL dataset audit | | AT | API Timeout | Tool call times out; agent never gets the result | Increase per-turn timeout |
bash# Find the result for the task_id in the latest run RESULTS=~/.cache/ruflo/gaia/results-latest.json node -e " const r = JSON.parse(require('fs').readFileSync('$RESULTS')); const q = r.results.find(x => x.task_id === '$TASK_ID'); console.log(JSON.stringify(q, null, 2)); "
Look at the trace output:
bashnode v3/@claude-flow/cli/bin/cli.js gaia-bench run \ --level 1 --limit 1 \ --task-id $TASK_ID \ --models claude-sonnet-4-6 \ --max-turns 20 \ --output json
| Failure | Action | |---------|--------| | TG — missing web_browse | Verify gaia-tools/index.ts exports web_browse; check tool registration | | TG — missing image OCR | Add image_describe tool call; verify GOOGLE_AI_API_KEY | | RM — reasoning | Add a system prompt instruction: "Before answering, list all facts you have gathered" | | EB — extraction | Test the FINAL_ANSWER_RE regex against the trace manually | | LI — loop | Add a tool-call deduplication guard in gaia-agent.ts | | AT — timeout | Set DEFAULT_PER_TURN_TIMEOUT_MS higher or use --max-turns flag |
bash# Re-run the single question node … gaia-bench run --task-id $TASK_ID --models $MODEL --output json # If now passing, store the pattern npx @claude-flow/cli@latest memory store \ --namespace gaia-debug-patterns \ --key "fix-$FAILURE_CODE-$(date +%Y%m%d)" \ --value "task_id=$TASK_ID, mode=$FAILURE_CODE, fix=$FIX_DESCRIPTION"
bashnode -e " const { createDefaultToolCatalogue } = require('./v3/@claude-flow/cli/src/benchmarks/gaia-tools/index.js'); const cat = createDefaultToolCatalogue({}); console.log('Tools registered:', cat.definitions.map(t => t.name)); "
Expected: web_search, file_read, web_browse, image_describe, python_exec
After resolving a debugging session, store the finding:
bashnpx @claude-flow/cli@latest memory store \ --namespace gaia-debug-patterns \ --key "session-$(date +%Y%m%d-%H%M)" \ --value '{"task_id":"$TASK_ID","failure_mode":"$CODE","fix":"$FIX","verified":true}'
Search for similar past failures:
bashnpx @claude-flow/cli@latest memory search \ --namespace gaia-debug-patterns \ --query "extraction bug final answer regex"
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-08 | fail→pass | 9,352 | 5,997 | -36% | 1 | 1 | 0% | 1,682 | 2,053 | +22% | 0 | 0 | — |
case-06 | fail→pass | 10,580 | 6,025 | -43% | 1 | 1 | 0% | 1,865 | 2,105 | +13% | 0 | 0 | — |
case-07 | fail→pass | 12,966 | 5,552 | -57% | 1 | 1 | 0% | 2,114 | 2,087 | -1% | 0 | 0 | — |
case-01 | fail→pass | 5,126 | 9,301 | +81% | 1 | 1 | 0% | 329 | 2,871 | +773% | 0 | 0 | — |
case-02 | fail→pass | 4,980 | 6,677 | +34% | 1 | 1 | 0% | 219 | 2,322 | +960% | 0 | 0 | — |
case-03 | fail→pass | 17,482 | 7,623 | -56% | 1 | 1 | 0% | 3,174 | 2,655 | -16% | 0 | 0 | — |
case-04 | fail→pass | 14,234 | 5,844 | -59% | 1 | 1 | 0% | 2,309 | 2,217 | -4% | 0 | 0 | — |
case-05 | fail→pass | 10,648 | 3,894 | -63% | 1 | 1 | 0% | 1,779 | 1,825 | +3% | 0 | 0 | — |
case-09 | fail→pass | 11,958 | 1,952 | -84% | 1 | 1 | 0% | 2,108 | 1,469 | -30% | 0 | 0 | — |
case-10 | fail→pass | 8,300 | 1,968 | -76% | 1 | 1 | 0% | 1,443 | 1,492 | +3% | 0 | 0 | — |
case-11 | fail→pass | 11,947 | 1,888 | -84% | 1 | 1 | 0% | 1,943 | 1,353 | -30% | 0 | 0 | — |
case-12 | fail→pass | 7,984 | 2,564 | -68% | 1 | 1 | 0% | 1,364 | 1,514 | +11% | 0 | 0 | — |
case-13 | fail→pass | 11,384 | 2,440 | -79% | 1 | 1 | 0% | 1,823 | 1,603 | -12% | 0 | 0 | — |
case-14 | fail→pass | 10,355 | 4,357 | -58% | 1 | 1 | 0% | 1,659 | 1,872 | +13% | 0 | 0 | — |
case-15 | fail→pass | 10,559 | 2,441 | -77% | 1 | 1 | 0% | 1,735 | 1,602 | -8% | 0 | 0 | — |
case-16 | fail→pass | 8,582 | 3,361 | -61% | 1 | 1 | 0% | 1,370 | 1,728 | +26% | 0 | 0 | — |
case-17 | fail→pass | 15,072 | 3,139 | -79% | 1 | 1 | 0% | 2,820 | 1,781 | -37% | 0 | 0 | — |
case-18 | fail→pass | 10,278 | 3,125 | -70% | 1 | 1 | 0% | 1,569 | 1,677 | +7% | 0 | 0 | — |
case-19 | fail→pass | 10,226 | 2,734 | -73% | 1 | 1 | 0% | 1,572 | 1,531 | -3% | 0 | 0 | — |
case-20 | pass→pass | 15,211 | 9,852 | -35% | 1 | 1 | 0% | 2,405 | 2,669 | +11% | 0 | 0 | — |
case-21 | pass→pass | 13,433 | 10,303 | -23% | 1 | 1 | 0% | 2,652 | 3,069 | +16% | 0 | 0 | — |
case-22 | pass→pass | 9,218 | 5,977 | -35% | 1 | 1 | 0% | 1,854 | 2,280 | +23% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 20 counted toward the lift figure. The other 2 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +86 percentage points is the difference between those two pass rates over the 20 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.