Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Walk through a complete GAIA benchmark→submit flow — from key resolution through HAL-compatible package generation
.claude/skills/ruvnet-gaia-submission/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-08 | ✗→✓ | ▲ Improved | -16% | 0% |
| case-10 | ✗→✓ | ▲ Improved | -33% | 0% |
| case-12 | ✗→✓ | ▲ Improved | -10% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 23% | 0% |
| case-14 | ✗→✓ | ▲ Improved | 122% | 0% |
Walk Claude Code through every step needed to go from a clean environment to a signed, HAL-compatible submission package ready to upload to the Princeton GAIA leaderboard.
When the user wants to:
Before starting, confirm these are available:
| Requirement | Check | |-------------|-------| | ANTHROPIC_API_KEY | echo ${ANTHROPIC_API_KEY:0:8}… (should show sk-ant-…) | | HF_TOKEN | echo ${HF_TOKEN:0:5}… (should show hf_…) | | Node.js 20+ | node --version | | CLI built | node v3/@claude-flow/cli/bin/cli.js --version |
bash# Run all pre-flight checks /gaia validate
If any check fails, resolve it before continuing.
Ask the user for their configuration:
claude-sonnet-4-6)bash/gaia cost --level=$LEVEL --limit=$LIMIT --models=$MODELS --voting=$VOTING
If projected cost > $5, show the estimate and ask: "This run will cost approximately $X. Proceed? (y/N)"
bash/gaia run --level=$LEVEL --limit=$LIMIT --models=$MODELS --voting=$VOTING
While running, progress is reported every 5 questions:
[12/53] 22.7% (5 passed of 22 scored) — est. remaining: $0.18Store the run summary in memory for history tracking:
bashnpx @claude-flow/cli@latest memory store \ --namespace gaia-runs \ --key "run-$(date +%Y%m%d-%H%M)" \ --value '{"level":$LEVEL,"model":"$MODEL","total":$TOTAL,"passed":$PASSED,"pass_rate":$RATE,"est_cost_usd":$COST}'
bash/gaia submit --results=~/.cache/ruflo/gaia/results-latest.json
This produces:
submission-<date>-<sha>/
├── results.jsonl ← HAL-compatible, one JSON per line
├── trajectories.jsonl ← full agent traces
├── metadata.json ← harness info, model, tool catalogue
├── audit-report.json ← ADR-167 pre-submission exploit-audit report
├── manifest.md.json ← Ed25519-signed witness (signs audit-report.json's hash)
└── README.md ← human summary + leaderboard comparisonPost-RDI (UC Berkeley broke 8 agent benchmarks — GAIA to ~98% — without solving a task), a signature alone is not enough: it proves the bytes are untampered, not that the score was earned. /gaia submit therefore runs a deterministic, $0 exploit audit before signing and refuses to build the leaderboard package on a CRITICAL failure unless --allow-dirty is passed. The audit report is signed into the witness manifest as an ADR-103 fix marker, so a ruflo GAIA submission attests both transport-integrity and earning-integrity.
If the gate blocks, treat it as a real finding — inspect audit-report.json (answer-leakage, no-work pass, oracle leakage, grader monkey-patching, an answer-key read outside the dataset dir, or dynamic eval/exec of task content in the runner) rather than reaching for --allow-dirty. The static source-scan family (answer-key-reads, dynamic-eval, judge-injection) enforces today with no trajectory instrumentation; the trajectory-fed checks the current schema cannot feed are reported as harness_gaps (ADR-167 §7), not passes.
bash/gaia leaderboard --level=$LEVEL /gaia history
Interpret the gap between ruflo's score and the leaderboard top-10. Identify the primary failure mode (tool gap, reasoning miss, extraction bug) using the /gaia-debugging skill if needed.
bashnpx @claude-flow/cli@latest hooks post-task \ --task-id "gaia-submission-$(date +%Y%m%d)" \ --success true \ --train-neural true
Store any discovered patterns:
bashnpx @claude-flow/cli@latest memory store \ --namespace gaia-patterns \ --key "submission-notes-$(date +%Y%m%d)" \ --value "Level $LEVEL, $MODEL: $NOTES"
This skill is intentionally structured to be benchmark-agnostic. The phase headers (validate → estimate → run → package → compare → learn) apply to SWE-bench, WebArena, and HumanEval with only phase 3-4 details changing.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 14,439 | 3,052 | -79% | 1 | 1 | 0% | 3,579 | 1,819 | -49% | 0 | 0 | — |
case-02 | fail→fail | 7,701 | 6,694 | -13% | 1 | 1 | 0% | 1,447 | 1,712 | +18% | 0 | 0 | — |
case-03 | fail→fail | 8,926 | 2,817 | -68% | 1 | 1 | 0% | 1,891 | 1,724 | -9% | 0 | 0 | — |
case-04 | pass→pass | 11,037 | 6,821 | -38% | 1 | 1 | 0% | 2,548 | 2,685 | +5% | 0 | 0 | — |
case-05 | pass→fail | 16,024 | 12,953 | -19% | 1 | 1 | 0% | 3,643 | 4,207 | +15% | 0 | 0 | — |
case-06 | fail→fail | 10,700 | 10,761 | +1% | 1 | 1 | 0% | 2,105 | 3,601 | +71% | 0 | 0 | — |
case-07 | fail→fail | 8,544 | 4,401 | -48% | 1 | 1 | 0% | 1,946 | 1,927 | -1% | 0 | 0 | — |
case-08 | fail→pass | 10,253 | 1,704 | -83% | 1 | 1 | 0% | 1,900 | 1,605 | -16% | 0 | 0 | — |
case-09 | pass→pass | 7,832 | 3,087 | -61% | 1 | 1 | 0% | 1,554 | 1,924 | +24% | 0 | 0 | — |
case-10 | fail→pass | 25,757 | 1,512 | -94% | 1 | 1 | 0% | 2,403 | 1,611 | -33% | 0 | 0 | — |
case-11 | fail→fail | 9,405 | 1,786 | -81% | 1 | 1 | 0% | 1,874 | 1,624 | -13% | 0 | 0 | — |
case-12 | fail→pass | 10,680 | 3,149 | -71% | 1 | 1 | 0% | 2,290 | 2,064 | -10% | 0 | 0 | — |
case-13 | fail→pass | 12,108 | 3,809 | -69% | 1 | 1 | 0% | 1,700 | 2,093 | +23% | 0 | 0 | — |
case-14 | fail→pass | 47,917 | 10,182 | -79% | 1 | 1 | 0% | 865 | 1,921 | +122% | 0 | 0 | — |
case-15 | fail→pass | 9,183 | 2,077 | -77% | 1 | 1 | 0% | 1,806 | 1,699 | -6% | 0 | 0 | — |
case-16 | fail→pass | 13,210 | 3,376 | -74% | 1 | 1 | 0% | 2,051 | 1,804 | -12% | 0 | 0 | — |
case-17 | fail→fail | 8,872 | 4,061 | -54% | 1 | 1 | 0% | 1,867 | 1,835 | -2% | 0 | 0 | — |
case-18 | fail→pass | 11,947 | 3,038 | -75% | 1 | 1 | 0% | 1,909 | 1,824 | -4% | 0 | 0 | — |
case-19 | fail→pass | 6,120 | 1,844 | -70% | 1 | 1 | 0% | 1,111 | 1,668 | +50% | 0 | 0 | — |
case-20 | fail→pass | 11,436 | 2,479 | -78% | 1 | 1 | 0% | 1,875 | 1,799 | -4% | 0 | 0 | — |
case-21 | pass→pass | 13,273 | 2,332 | -82% | 1 | 1 | 0% | 2,429 | 1,749 | -28% | 0 | 0 | — |
case-22 | fail→pass | 8,557 | 1,299 | -85% | 1 | 1 | 0% | 1,744 | 1,482 | -15% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 20 counted toward the lift figure. The other 2 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +40 percentage points is the difference between those two pass rates over the 20 comparable cases. 3 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.