Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Measure declared project fitness goals
.claude/skills/boshu2-fitness/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | -61% | 0% |
| case-05 | ✗→✓ | ▲ Improved | -70% | 0% |
| case-06 | ✗→✓ | ▲ Improved | -46% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -58% | 0% |
| case-08 | ✗→✓ | ▲ Improved | -67% | 0% |
Inspect the active goals document and run only the caller-selected measurement, validation, drift, history, export, or meta-goal command.
Renamed from goals (2026-07-29): the semantic skill is fitness; the ao goals CLI command family is a separate product surface and keeps its name. A thin goals compatibility alias resolves to this skill.
Measurement stays trustworthy only because it cannot mutate what it measures; the moment a fitness report edits a goal, the next report measures the editor, not the project.
Named failure mode — advice creep: a measurement report that ends with "you should…" has silently become work selection.
Anti-pattern: padding the report with recommendations to look helpful. Corrective: return the numbers, the evidence gaps, and checked/not-checked scope, and let the caller decide.
GOALS.md when both Markdown and legacy YAML exist.mutate goals.
measure,drift, and export persist a best-effort JSON snapshot under the fixed derived path .agents/ao/goals/baselines/. render --out <file> writes a Gherkin spec to whatever path the caller names — the CLI does not constrain it, so never point --out at the goals source or any non-derived file.
All eight subcommands read the goals source without mutating it. The snapshot write (fixed derived path) and the render --out write (caller-chosen path, caller's responsibility) are the only side effects.
bashao goals measure --json ao goals validate --json ao goals drift ao goals history ao goals export ao goals meta --json ao goals scenarios ao goals render # append --out <file> to write the spec instead of stdout
Run the requested command once. Return the command, exit code, goal-level results, aggregate measurement, missing evidence, and checked/not-checked scope. Then stop.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 10,033 | 4,894 | -51% | 1 | 1 | 0% | 1,879 | 749 | -60% | 0 | 0 | — |
case-02 | fail→fail | 9,170 | 6,029 | -34% | 1 | 1 | 0% | 1,388 | 803 | -42% | 0 | 0 | — |
case-03 | fail→fail | 7,102 | 8,037 | +13% | 1 | 1 | 0% | 1,119 | 1,198 | +7% | 0 | 0 | — |
case-04 | fail→pass | 12,793 | 2,105 | -84% | 1 | 1 | 0% | 2,118 | 831 | -61% | 0 | 0 | — |
case-05 | fail→pass | 15,581 | 1,982 | -87% | 1 | 1 | 0% | 2,721 | 824 | -70% | 0 | 0 | — |
case-06 | fail→pass | 9,138 | 1,691 | -81% | 1 | 1 | 0% | 1,424 | 767 | -46% | 0 | 0 | — |
case-07 | fail→pass | 12,466 | 2,393 | -81% | 1 | 1 | 0% | 1,976 | 832 | -58% | 0 | 0 | — |
case-08 | fail→pass | 15,536 | 2,462 | -84% | 1 | 1 | 0% | 2,592 | 856 | -67% | 0 | 0 | — |
case-09 | fail→pass | 9,777 | 4,381 | -55% | 1 | 1 | 0% | 1,659 | 1,112 | -33% | 0 | 0 | — |
case-10 | fail→pass | 18,447 | 1,734 | -91% | 1 | 1 | 0% | 2,989 | 709 | -76% | 0 | 0 | — |
case-11 | pass→pass | 4,319 | 2,051 | -53% | 1 | 1 | 0% | 713 | 757 | +6% | 0 | 0 | — |
case-12 | pass→fail | 10,739 | 2,043 | -81% | 1 | 1 | 0% | 1,820 | 723 | -60% | 0 | 0 | — |
case-13 | fail→pass | 8,022 | 5,342 | -33% | 1 | 1 | 0% | 1,259 | 1,364 | +8% | 0 | 0 | — |
case-14 | fail→fail | 5,442 | 1,446 | -73% | 1 | 1 | 0% | 829 | 745 | -10% | 0 | 0 | — |
case-15 | fail→fail | 9,759 | 4,741 | -51% | 1 | 1 | 0% | 1,604 | 1,272 | -21% | 0 | 0 | — |
case-16 | fail→pass | 17,152 | 13,082 | -24% | 1 | 1 | 0% | 2,577 | 2,024 | -21% | 0 | 0 | — |
case-17 | fail→pass | 18,287 | 1,520 | -92% | 1 | 1 | 0% | 2,898 | 714 | -75% | 0 | 0 | — |
case-18 | pass→pass | 6,826 | 3,452 | -49% | 1 | 1 | 0% | 1,105 | 1,021 | -8% | 0 | 0 | — |
case-19 | fail→pass | 14,737 | 1,473 | -90% | 1 | 1 | 0% | 2,142 | 767 | -64% | 0 | 0 | — |
case-20 | fail→pass | 3,340 | 3,255 | -3% | 1 | 1 | 0% | 211 | 971 | +360% | 0 | 0 | — |
case-21 | fail→pass | 13,985 | 10,937 | -22% | 1 | 1 | 0% | 2,188 | 2,306 | +5% | 0 | 0 | — |
case-22 | fail→pass | 6,692 | 3,496 | -48% | 1 | 1 | 0% | 1,023 | 1,022 | -0% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 19 counted toward the lift figure. The other 3 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +59 percentage points is the difference between those two pass rates over the 19 comparable cases. 3 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.