Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Delegate bounded work to other AI agents while preserving context, ownership, and progress checks.
.claude/skills/sickn33-delegating-to-agents/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 49% | 0% |
| case-04 | ✗→✓ | ▲ Improved | -39% | 0% |
| case-05 | ✗→✓ | ▲ Improved | -19% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 8% | 0% |
| case-08 | ✗→✓ | ▲ Improved | -18% | 0% |
pi in a cmux terminal). All Pi agents run opus-4.8-fast via OpenRouter at xhigh reasoning effort.read /tmp/task.md and follow it.cmux send --surface surface:N "your prompt". The recurring bug is emitting \" — in bash that's literal-broken and dies with unexpected EOF. Inside the prompt, avoid apostrophes and literal double quotes (write "dont", "wont", "lets"); rephrase instead of escaping. If a send failed, the cause was the escaped \", not the quote type.cmux send --surface surface:N then cmux send-key --surface surface:N enter. There is NO send-surface or send-key-surface.Keep sleeps SHORT: start at 3-5s, re-check, repeat. Don't sleep 30. Pi and Hermes (opus-4.8-fast) launch and respond within seconds; scale up only for genuinely heavy tasks. After every check, send the user a one-line status: what the agent is doing and whether it's on track.
Claude Code note: after it finishes, it may prefill a predicted next user message — that draft is Claude, not the user.
SSH in first and launch the agent ON the VPS (e.g. codex --yolo), then drive that on-box agent. Don't run an agent locally and have it SSH for every step.
All four use the portable SKILL.md standard; project skills win over global.
~/.pi/agent/skills/.codex exec for CI; reads AGENTS.md. Skills: ~/.codex/skills/..claude/ conventions, live skill hot-reload. Skills: ~/.claude/skills/.~/.hermes/skills/.pty=true.claude --print --permission-mode bypassPermissions (no PTY).User request:
> Delegate these independent work items with complete context, clear ownership, and progress checks.
davidondrej/skills; verify local paths, tools, credentials, and agent features before acting.| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 13,077 | 8,161 | -38% | 1 | 1 | 0% | 1,638 | 2,433 | +49% | 0 | 0 | — |
case-02 | fail→fail | 8,388 | 7,841 | -7% | 1 | 1 | 0% | 1,357 | 1,560 | +15% | 0 | 0 | — |
case-03 | fail→fail | 10,538 | 9,038 | -14% | 1 | 1 | 0% | 1,885 | 2,831 | +50% | 0 | 0 | — |
case-04 | fail→pass | 13,573 | 3,128 | -77% | 1 | 1 | 0% | 2,379 | 1,443 | -39% | 0 | 0 | — |
case-05 | fail→pass | 10,982 | 3,284 | -70% | 1 | 1 | 0% | 2,002 | 1,617 | -19% | 0 | 0 | — |
case-06 | pass→pass | 13,110 | 4,829 | -63% | 1 | 1 | 0% | 2,105 | 1,890 | -10% | 0 | 0 | — |
case-07 | fail→pass | 7,568 | 3,755 | -50% | 1 | 1 | 0% | 1,475 | 1,586 | +8% | 0 | 0 | — |
case-08 | fail→pass | 9,223 | 2,418 | -74% | 1 | 1 | 0% | 1,693 | 1,395 | -18% | 0 | 0 | — |
case-09 | fail→pass | 11,306 | 2,290 | -80% | 1 | 1 | 0% | 2,069 | 1,292 | -38% | 0 | 0 | — |
case-10 | fail→pass | 7,238 | 2,291 | -68% | 1 | 1 | 0% | 1,385 | 1,332 | -4% | 0 | 0 | — |
case-11 | fail→pass | 8,655 | 4,342 | -50% | 1 | 1 | 0% | 1,641 | 1,824 | +11% | 0 | 0 | — |
case-12 | fail→pass | 10,411 | 1,722 | -83% | 1 | 1 | 0% | 2,075 | 1,246 | -40% | 0 | 0 | — |
case-13 | pass→pass | 3,913 | 1,563 | -60% | 1 | 1 | 0% | 712 | 1,189 | +67% | 0 | 0 | — |
case-14 | fail→pass | 6,281 | 1,398 | -78% | 1 | 1 | 0% | 1,101 | 1,150 | +4% | 0 | 0 | — |
case-15 | fail→pass | 7,961 | 1,357 | -83% | 1 | 1 | 0% | 1,419 | 1,091 | -23% | 0 | 0 | — |
case-16 | pass→pass | 7,479 | 2,460 | -67% | 1 | 1 | 0% | 1,452 | 1,349 | -7% | 0 | 0 | — |
case-17 | fail→pass | 6,743 | 2,272 | -66% | 1 | 1 | 0% | 1,183 | 1,242 | +5% | 0 | 0 | — |
case-18 | fail→pass | 11,729 | 2,772 | -76% | 1 | 1 | 0% | 1,942 | 1,473 | -24% | 0 | 0 | — |
case-19 | pass→pass | 10,834 | 3,455 | -68% | 1 | 1 | 0% | 1,825 | 1,620 | -11% | 0 | 0 | — |
case-20 | pass→pass | 6,474 | 4,276 | -34% | 1 | 1 | 0% | 1,335 | 1,813 | +36% | 0 | 0 | — |
case-21 | pass→pass | 8,986 | 4,959 | -45% | 1 | 1 | 0% | 1,691 | 1,867 | +10% | 0 | 0 | — |
case-22 | pass→pass | 14,216 | 7,584 | -47% | 1 | 1 | 0% | 2,530 | 2,343 | -7% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +59 percentage points is the difference between those two pass rates over the 22 comparable cases.
The publisher has shipped newer versions since this run, so these numbers describe v1, not the version currently listed.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.