Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Orchestrate task execution via fresh subagents with dispatch, monitoring, and result collection
.claude/skills/yogsoth-ai-subagent-execution-loop/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | 40% | 0% |
| case-05 | ✗→✓ | ▲ Improved | -26% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 7% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -35% | 0% |
| case-08 | ✗→✓ | ▲ Improved | -36% | 0% |
执行循环直接 load superpowers:subagent-driven-development 当引擎,不再自写循环伪码。 每个任务在引擎内部按以下贯穿规则执行:
Skill load ponytail:ponytail —— 任务开始即开启精简反射。Skill load superpowers:test-driven-development —— 每任务 RED → GREEN → REFACTOR。Skill load superpowers:requesting-code-review —— 任务 diff 出来后派 reviewer 子代理。Skill load ponytail:ponytail-review —— code-review 之后,对 diff 再过一道过度工程审(delete/stdlib/native/yagni/shrink)。
Skill load superpowers:receiving-code-review —— 核验 review 反馈再落地,反馈错就反驳。DARE 原生的 implementer-dispatch / execution-monitoring / result-collection 作为 引擎内每任务的派单、盯状态、收结果三步保留。
| Condition | Action | |-----------|--------| | model 选择 / retry / 超时 / 死锁 | 交给 superpowers:subagent-driven-development 引擎处理 | | code-review 后 | 先 ponytail:ponytail-review 查过度工程,再 receiving-code-review 落地 | | Budget < 10% remaining | 停,报告 partial(DARE 预算治理) | | >50% 关键路径 BLOCKED | 中止执行(DARE 排程判据) |
<!-- BEGIN available-tables (generated) --> <!-- external rows hand-maintained; do not regenerate this file -->
Optional, no fixed order; the final leaf is always a sop.
| SOP | When to use | | --- | --- | | execution-monitoring | Monitor execution progress, detect anomalies, and report status | | implementer-dispatch | Dispatch execution subagent — select model by complexity, construct prompt with full task context | | ponytail:ponytail | Lazy-senior reflex: simplest thing that holds; mark every deliberate shortcut | | ponytail:ponytail-review | Audit the diff for over-engineering (delete/stdlib/native/yagni/shrink) | | result-collection | Collect experiment outputs — metrics, logs, artifacts — into structured result set | | superpowers:receiving-code-review | Verify review feedback before applying; push back when wrong | | superpowers:requesting-code-review | Dispatch a code-reviewer subagent after each task | | superpowers:subagent-driven-development | Execute the plan via a fresh subagent per task with two-stage review | | superpowers:test-driven-development | RED -> GREEN -> REFACTOR per task |
<!-- END available-tables (generated) -->
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 11,407 | 17,130 | +50% | 1 | 1 | 0% | 999 | 1,327 | +33% | 0 | 0 | — |
case-02 | fail→fail | 22,703 | 39,244 | +73% | 1 | 1 | 0% | 3,533 | 1,237 | -65% | 0 | 0 | — |
case-03 | fail→fail | 14,092 | 15,715 | +12% | 1 | 1 | 0% | 2,511 | 1,062 | -58% | 0 | 0 | — |
case-04 | fail→pass | 51,315 | 11,466 | -78% | 1 | 1 | 0% | 1,365 | 1,914 | +40% | 0 | 0 | — |
case-05 | fail→pass | 19,065 | 5,564 | -71% | 1 | 1 | 0% | 2,181 | 1,607 | -26% | 0 | 0 | — |
case-06 | fail→pass | 27,859 | 8,943 | -68% | 1 | 1 | 0% | 1,273 | 1,366 | +7% | 0 | 0 | — |
case-07 | fail→pass | 16,212 | 11,962 | -26% | 1 | 1 | 0% | 2,426 | 1,588 | -35% | 0 | 0 | — |
case-08 | fail→pass | 23,594 | 11,905 | -50% | 1 | 1 | 0% | 2,915 | 1,865 | -36% | 0 | 0 | — |
case-09 | pass→pass | 17,043 | 13,958 | -18% | 1 | 1 | 0% | 1,941 | 2,090 | +8% | 0 | 0 | — |
case-10 | pass→pass | 11,057 | 8,601 | -22% | 1 | 1 | 0% | 1,737 | 1,270 | -27% | 0 | 0 | — |
case-11 | fail→pass | 17,304 | 8,284 | -52% | 1 | 1 | 0% | 1,830 | 1,276 | -30% | 0 | 0 | — |
case-12 | fail→pass | 15,912 | 9,914 | -38% | 1 | 1 | 0% | 1,684 | 1,560 | -7% | 0 | 0 | — |
case-13 | fail→fail | 12,973 | 8,882 | -32% | 1 | 1 | 0% | 1,215 | 1,412 | +16% | 0 | 0 | — |
case-14 | fail→fail | 17,087 | 2,941 | -83% | 1 | 1 | 0% | 1,879 | 1,185 | -37% | 0 | 0 | — |
case-15 | pass→pass | 16,407 | 4,500 | -73% | 1 | 1 | 0% | 1,717 | 1,370 | -20% | 0 | 0 | — |
case-16 | pass→pass | 15,912 | 10,269 | -35% | 1 | 1 | 0% | 1,701 | 1,663 | -2% | 0 | 0 | — |
case-17 | pass→pass | 7,400 | 2,180 | -71% | 1 | 1 | 0% | 1,151 | 1,083 | -6% | 0 | 0 | — |
case-18 | pass→pass | 16,137 | 7,464 | -54% | 1 | 1 | 0% | 1,804 | 1,141 | -37% | 0 | 0 | — |
case-19 | fail→pass | 15,791 | 6,820 | -57% | 1 | 1 | 0% | 1,711 | 956 | -44% | 0 | 0 | — |
case-20 | pass→pass | 14,127 | 15,405 | +9% | 1 | 1 | 0% | 2,297 | 3,278 | +43% | 0 | 0 | — |
case-21 | pass→pass | 22,510 | 24,836 | +10% | 1 | 1 | 0% | 2,664 | 4,596 | +73% | 0 | 0 | — |
case-22 | pass→pass | 24,081 | 29,371 | +22% | 1 | 1 | 0% | 4,247 | 6,095 | +44% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 18 counted toward the lift figure. The other 4 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +36 percentage points is the difference between those two pass rates over the 18 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.