Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Report observable AgentOps evidence without
.claude/skills/boshu2-status/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 17% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 91% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 109% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 96% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 15% | 0% |
A status snapshot is trustworthy exactly when every line traces to an artifact that exists on disk right now; the first inferred line turns the report into a guess wearing a report's clothes.
Report only observable local facts: available intent and verdict artifacts and their counts; deterministic check results; evidence recency; and unavailable or corrupt sources. The canonical durable stores are .agents/ao/intents/sha256 and .agents/ao/verdicts/sha256. Subject manifests are caller-supplied: report them only when the caller names their location, otherwise disclose them as not_checked. When .agents/ao evidence exists, report which stored artifact kind is newest and label that conclusion as evidence recency, not runtime phase or process activity. The ao status snapshot emits the two counts plus that newest-artifact recency conclusion, not a per-artifact digest or timestamp listing; report a specific artifact's digest or timestamp only when the caller asks about a named artifact.
Always disclose checked and not_checked. Runtime phase, execution elapsed time, tool-call activity, and remaining work are not_checked unless a caller provides a separate authoritative source for them.
ao status is the evidence-store view. It validates content-addressed artifact names and content before counting them, reports corrupt and unavailable entries, and shows only intent/verdict counts plus evidence recency. Retired legacy surfaces (session indexes, provenance summaries, knowledge-health signals) are not aggregated into this command.
Distinct state classes stay distinct in every status report: tracker state belongs to the tracker, Git state to Git, factory or runtime state to the selected factory's own doors, deterministic-check results to the executable that produced them, and semantic validation to a fresh verdict. Report each from its own authority, never merged into one blended "health" claim — factory-complete, checks-green, and AgentOps-PASS are three different facts.
Status does not inspect work queues, assign priority, claim work, infer a next action, repair records, govern retries, or change any state. Optional Git or tracker metadata may be displayed only when the caller supplies it; absence cannot change the report interpretation.
Named failure mode — recency-as-activity: reading "newest artifact is a verdict" as "validation is running", which invents a runtime phase from a timestamp.
Anti-pattern: filling not_checked gaps with plausible narrative so the snapshot feels complete. Corrective: report the gap as a gap; an honest hole outranks a smooth story.
Return the snapshot and stop.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 11,048 | 14,932 | +35% | 1 | 1 | 0% | 1,755 | 2,050 | +17% | 0 | 0 | — |
case-02 | fail→pass | 6,023 | 7,824 | +30% | 1 | 1 | 0% | 1,088 | 2,080 | +91% | 0 | 0 | — |
case-03 | fail→pass | 10,011 | 22,586 | +126% | 1 | 1 | 0% | 1,714 | 3,582 | +109% | 0 | 0 | — |
case-04 | fail→pass | 4,616 | 4,579 | -1% | 1 | 1 | 0% | 754 | 1,476 | +96% | 0 | 0 | — |
case-05 | fail→pass | 7,303 | 5,201 | -29% | 1 | 1 | 0% | 1,276 | 1,473 | +15% | 0 | 0 | — |
case-06 | fail→pass | 11,216 | 9,495 | -15% | 1 | 1 | 0% | 1,836 | 2,428 | +32% | 0 | 0 | — |
case-07 | fail→pass | 11,826 | 4,344 | -63% | 1 | 1 | 0% | 1,847 | 1,293 | -30% | 0 | 0 | — |
case-08 | fail→pass | 58,695 | 12,823 | -78% | 1 | 1 | 0% | 1,810 | 3,020 | +67% | 0 | 0 | — |
case-09 | fail→pass | 10,971 | 4,304 | -61% | 1 | 1 | 0% | 2,202 | 1,354 | -39% | 0 | 0 | — |
case-10 | fail→pass | 2,000 | 5,392 | +170% | 1 | 1 | 0% | 318 | 1,526 | +380% | 0 | 0 | — |
case-11 | fail→pass | 5,087 | 8,508 | +67% | 1 | 1 | 0% | 788 | 2,035 | +158% | 0 | 0 | — |
case-12 | fail→pass | 8,101 | 3,803 | -53% | 1 | 1 | 0% | 1,349 | 1,206 | -11% | 0 | 0 | — |
case-13 | fail→fail | 7,996 | 8,939 | +12% | 1 | 1 | 0% | 1,261 | 2,125 | +69% | 0 | 0 | — |
case-18 | fail→pass | 6,975 | 7,819 | +12% | 1 | 1 | 0% | 1,091 | 2,033 | +86% | 0 | 0 | — |
case-14 | fail→pass | 5,452 | 6,447 | +18% | 1 | 1 | 0% | 845 | 1,770 | +109% | 0 | 0 | — |
case-15 | fail→pass | 7,259 | 8,366 | +15% | 1 | 1 | 0% | 1,237 | 2,172 | +76% | 0 | 0 | — |
case-16 | fail→pass | 13,013 | 2,989 | -77% | 1 | 1 | 0% | 1,957 | 1,067 | -45% | 0 | 0 | — |
case-17 | pass→pass | 7,851 | 11,876 | +51% | 1 | 1 | 0% | 1,199 | 1,669 | +39% | 0 | 0 | — |
case-19 | pass→pass | 12,000 | 4,484 | -63% | 1 | 1 | 0% | 1,940 | 1,295 | -33% | 0 | 0 | — |
case-20 | fail→pass | 7,366 | 4,365 | -41% | 1 | 1 | 0% | 1,223 | 1,281 | +5% | 0 | 0 | — |
case-21 | pass→pass | 6,663 | 3,081 | -54% | 1 | 1 | 0% | 990 | 992 | +0% | 0 | 0 | — |
case-22 | pass→pass | 3,162 | 2,391 | -24% | 1 | 1 | 0% | 440 | 914 | +108% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 21 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +77 percentage points is the difference between those two pass rates over the 21 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.