Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when managing, triaging, restarting, escalating, or summarizing Codewhale Agent Fleet runs and workers.
.claude/skills/hmbown-fleet-manager/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 1487% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 1243% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 47% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 54% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 55% | 0% |
Use this skill when acting as a manager agent for Codewhale fleet runs. Your job is to classify worker state, choose the narrowest safe typed action, and leave a ledgered receipt or a safe escalation draft.
codewhale fleet status,inspect, logs, artifacts, interrupt, restart, resume, stop (stop requires --all), and the Runtime API endpoints.
.codewhale/fleet.jsonl, host logs, or remote files directlyunless the typed command or API is missing required evidence.
user or run config explicitly authorizes sending. Draft the message instead.
oversized logs in a summary or escalation.
status output. If no worker is named, start with codewhale fleet status.
codewhale fleet inspect <worker-id> or thematching Runtime API worker endpoint (GET /v1/fleet/workers/{worker_id}).
codewhale fleet logs <worker-id> andcodewhale fleet artifacts <worker-id>, or the Runtime API equivalents (GET /v1/fleet/runs/{run_id}/receipts/{task_id}/evidence and GET /v1/fleet/runs/{run_id}/events/replay). Summarize artifact refs, not full payloads.
transient failure: transport error, timeout, stale heartbeat, hostunavailable, or retryable provider/network failure.
task failure: worker completed the task but the result is wrong,missing required artifacts, or reports a domain error.
verifier failure: scorer/verifier failed or disagrees with the workerresult.
needs-human: missing authority, unsafe secret boundary, destructiveaction, repeated restart exhaustion, ambiguous product decision, or conflict between artifacts and verifier.
codewhale fleet resume <run-id> (idempotent reconcile).
codewhale fleet restart <worker-id>.unless the task spec says retrying can produce new evidence.
verifier cannot be corrected through a typed action.
evidence commands, artifact refs, and next owner.
Restart only when all of these are true:
Escalate when any of these are true:
evidence,
Use this shape for Slack/PagerDuty drafts. Keep logs to three short lines or an artifact ref.
textCodewhale fleet needs attention Run: <run-id> Worker: <worker-id> Task: <task-id or unknown> Classification: <transient failure | task failure | verifier failure | needs-human> Reason: <one sentence, no secrets> Latest typed evidence: codewhale fleet inspect <worker-id>; codewhale fleet artifacts <worker-id> Safe log excerpt: <3 lines max or "see artifact <ref>"> Requested decision: <restart approval | verifier review | task owner review | permission decision>
End every Fleet Manager response with a compact receipt:
textFleet receipt Run: <run-id> Workers checked: <count/list> Classification: <state> Action: <restart/interrupt/stop/resume/escalation draft/no-op> Ledger expectation: <typed action should be recorded | draft only, no send> Artifacts reviewed: <refs> Follow-up owner: <manager | task owner | human>
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-19 | pass→pass | 8,412 | 12,160 | +45% | 1 | 1 | 0% | 1,159 | 1,916 | +65% | 0 | 0 | — |
case-01 | fail→pass | 16,856 | 30,541 | +81% | 1 | 1 | 0% | 223 | 3,539 | +1487% | 0 | 0 | — |
case-02 | fail→pass | 17,781 | 20,763 | +17% | 1 | 1 | 0% | 301 | 4,041 | +1243% | 0 | 0 | — |
case-13 | fail→pass | 17,583 | 15,255 | -13% | 1 | 1 | 0% | 1,916 | 2,825 | +47% | 0 | 0 | — |
case-03 | fail→fail | 17,391 | 12,819 | -26% | 1 | 1 | 0% | 364 | 1,394 | +283% | 0 | 0 | — |
case-04 | fail→fail | 47,074 | 21,325 | -55% | 1 | 1 | 0% | 1,309 | 4,194 | +220% | 0 | 0 | — |
case-05 | pass→fail | 21,704 | 22,319 | +3% | 1 | 1 | 0% | 3,477 | 3,642 | +5% | 0 | 0 | — |
case-06 | pass→fail | 25,611 | 25,710 | +0% | 1 | 1 | 0% | 3,799 | 4,591 | +21% | 0 | 0 | — |
case-07 | fail→pass | 9,667 | 7,067 | -27% | 1 | 1 | 0% | 1,488 | 2,294 | +54% | 0 | 0 | — |
case-08 | fail→pass | 14,708 | 10,906 | -26% | 1 | 1 | 0% | 1,365 | 2,113 | +55% | 0 | 0 | — |
case-09 | fail→pass | 33,561 | 13,676 | -59% | 1 | 1 | 0% | 1,967 | 2,494 | +27% | 0 | 0 | — |
case-10 | fail→pass | 13,904 | 15,089 | +9% | 1 | 1 | 0% | 1,128 | 2,680 | +138% | 0 | 0 | — |
case-11 | fail→fail | 20,357 | 10,757 | -47% | 1 | 1 | 0% | 1,932 | 2,858 | +48% | 0 | 0 | — |
case-12 | pass→pass | 9,512 | 12,791 | +34% | 1 | 1 | 0% | 1,464 | 2,180 | +49% | 0 | 0 | — |
case-14 | fail→pass | 13,792 | 12,011 | -13% | 1 | 1 | 0% | 1,504 | 1,708 | +14% | 0 | 0 | — |
case-15 | fail→pass | 13,387 | 13,424 | +0% | 1 | 1 | 0% | 1,142 | 2,464 | +116% | 0 | 0 | — |
case-16 | pass→pass | 5,550 | 15,062 | +171% | 1 | 1 | 0% | 788 | 2,830 | +259% | 0 | 0 | — |
case-17 | pass→pass | 15,232 | 11,484 | -25% | 1 | 1 | 0% | 1,527 | 2,055 | +35% | 0 | 0 | — |
case-18 | fail→pass | 13,892 | 8,451 | -39% | 1 | 1 | 0% | 1,941 | 2,384 | +23% | 0 | 0 | — |
case-20 | fail→pass | 16,896 | 10,718 | -37% | 1 | 1 | 0% | 1,640 | 1,820 | +11% | 0 | 0 | — |
case-21 | fail→pass | 18,196 | 5,038 | -72% | 1 | 1 | 0% | 2,026 | 1,474 | -27% | 0 | 0 | — |
case-22 | fail→pass | 12,213 | 6,815 | -44% | 1 | 1 | 0% | 1,651 | 2,065 | +25% | 0 | 0 | — |
case-23 | fail→pass | 12,838 | 10,001 | -22% | 1 | 1 | 0% | 1,179 | 1,902 | +61% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 19 counted toward the lift figure. The other 4 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +52 percentage points is the difference between those two pass rates over the 19 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
| Model | Method | Date | Lift |
|---|---|---|---|
| gemini-3.6-flash | verified | 9/8/2026 | +38% |
| gemini-3.6-flash | verified | 9/2/2026 | +17% |
| gemini-3.6-flash | verified | 8/7/2026 | +36% |
Other measured skills in the registry, with their headline benchmark lift.