Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Orchestrate Spec Kitty work-package implementation and review loops until all packages are done, approved, or correctly rejected.
.claude/skills/priivacy-ai-spk-run-implement-review/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | -34% | 0% |
| case-05 | ✗→✓ | ▲ Improved | -51% | 0% |
| case-06 | ✗→✓ | ▲ Improved | -29% | 0% |
| case-09 | ✗→✓ | ▲ Improved | -46% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 17% | 0% |
Use this skill when a mission is in implementation/review lanes or the user wants the agent to coordinate implementers and reviewers.
spk-run-next to identify the active WP and lane.spk-run-review-wp.implementation attempt.
move-task --to for_review runs the pre-review regression gate synchronously as part of the transition (#2570/#2493/#2555): it derives the affected test scope for the Mission's changed files and runs it at head before the command returns. That scoped run is a real subprocess and can take from seconds to a few minutes depending on scope size.
completion — treat the move-task --to for_review command as still in flight until it exits, not as a fire-and-forget hand-back. Do not assume control returns immediately, and do not kill or retry the command while it is still running.
--skip-pre-review-gate.To disable it process-wide, set SPEC_KITTY_SYNC_DISABLE or SPEC_KITTY_SYNC_MINIMAL_IMPORT (legacy-compatible gate opt-outs rather than a new third env var). Either opt-out skips the gate before it resolves a workspace or spawns the subprocess.
observing a sub-agent that appears to hang on a for_review transition should first check whether the gate's scoped test subprocess is still running before treating the sub-agent as stuck.
For detailed orchestration mechanics, use spec-kitty-implement-review when available.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 12,766 | 15,977 | +25% | 1 | 1 | 0% | 1,248 | 751 | -40% | 0 | 0 | — |
case-02 | fail→fail | 27,659 | 15,965 | -42% | 1 | 1 | 0% | 3,024 | 775 | -74% | 0 | 0 | — |
case-03 | fail→fail | 14,990 | 15,466 | +3% | 1 | 1 | 0% | 204 | 800 | +292% | 0 | 0 | — |
case-04 | fail→pass | 11,786 | 7,372 | -37% | 1 | 1 | 0% | 1,226 | 805 | -34% | 0 | 0 | — |
case-05 | fail→pass | 15,228 | 7,067 | -54% | 1 | 1 | 0% | 1,706 | 832 | -51% | 0 | 0 | — |
case-06 | fail→pass | 16,958 | 10,777 | -36% | 1 | 1 | 0% | 1,904 | 1,354 | -29% | 0 | 0 | — |
case-07 | pass→pass | 17,930 | 8,681 | -52% | 1 | 1 | 0% | 2,237 | 1,101 | -51% | 0 | 0 | — |
case-08 | pass→pass | 16,841 | 8,917 | -47% | 1 | 1 | 0% | 2,044 | 1,150 | -44% | 0 | 0 | — |
case-09 | fail→pass | 17,674 | 7,806 | -56% | 1 | 1 | 0% | 1,829 | 995 | -46% | 0 | 0 | — |
case-10 | fail→pass | 11,994 | 8,761 | -27% | 1 | 1 | 0% | 1,059 | 1,237 | +17% | 0 | 0 | — |
case-15 | fail→pass | 20,235 | 6,645 | -67% | 1 | 1 | 0% | 3,079 | 727 | -76% | 0 | 0 | — |
case-11 | fail→pass | 15,959 | 6,905 | -57% | 1 | 1 | 0% | 1,935 | 760 | -61% | 0 | 0 | — |
case-12 | fail→pass | 21,604 | 7,534 | -65% | 1 | 1 | 0% | 2,843 | 920 | -68% | 0 | 0 | — |
case-13 | fail→pass | 14,065 | 7,517 | -47% | 1 | 1 | 0% | 1,581 | 901 | -43% | 0 | 0 | — |
case-14 | fail→pass | 14,193 | 7,618 | -46% | 1 | 1 | 0% | 1,466 | 923 | -37% | 0 | 0 | — |
case-16 | pass→pass | 11,316 | 7,259 | -36% | 1 | 1 | 0% | 941 | 928 | -1% | 0 | 0 | — |
case-17 | pass→pass | 12,485 | 7,832 | -37% | 1 | 1 | 0% | 1,176 | 986 | -16% | 0 | 0 | — |
case-18 | pass→pass | 19,219 | 18,107 | -6% | 1 | 1 | 0% | 2,378 | 1,413 | -41% | 0 | 0 | — |
case-19 | fail→pass | 20,133 | 9,744 | -52% | 1 | 1 | 0% | 2,396 | 1,413 | -41% | 0 | 0 | — |
case-20 | fail→fail | 15,537 | 15,431 | -1% | 1 | 1 | 0% | 224 | 684 | +205% | 0 | 0 | — |
case-21 | pass→pass | 13,232 | 18,736 | +42% | 1 | 1 | 0% | 1,561 | 1,741 | +12% | 0 | 0 | — |
case-22 | pass→fail | 35,764 | 26,922 | -25% | 1 | 1 | 0% | 2,391 | 1,198 | -50% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 17 counted toward the lift figure. The other 5 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +45 percentage points is the difference between those two pass rates over the 17 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
| Model | Method | Date | Lift |
|---|---|---|---|
| gemini-3.6-flash | verified | 8/13/2026 | +56% |
Other measured skills in the registry, with their headline benchmark lift.