Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Set a Spec Kitty work-package review verdict through the deterministic event-log seam, so every agent harness records approve/reject the same way.
.claude/skills/priivacy-ai-spk-run-verdict-capture/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 8% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 270% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 144% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 0% | 0% |
| case-10 | ✗→✓ | ▲ Improved | -7% | 0% |
Use this skill whenever an agent (any of the 13 slash-command harnesses or the shared Agent-Skill harnesses) needs to record the outcome of a WP review — approve, reject, or otherwise. It documents the single deterministic seam so verdicts never diverge by harness.
The sole authority for a WP's review verdict is the review_result event in the mission's status.events.jsonl (read back via specify_cli.status.event_sourced_review_result). The review-cycle-N.md render is non-authoritative prose — do not hand-edit it to change a verdict, and never treat it as the source of truth. Because the verdict is an event, it is append-only, deterministic, and identical across every agent.
bash spec-kitty agent tasks move-task <WP> --to approved --mission <slug> --note "<why>"
bash spec-kitty agent tasks move-task <WP> --to planned \ --review-feedback-file <path/to/review-feedback-N.md> --mission <slug>
bash spec-kitty agent status emit <WP> --to done --actor <name> \ --evidence-json '{"review": {"reviewer": "<name>", "verdict": "approved", "reference": "<PR/commit>"}}'
Each command emits the review_result (or review-bearing transition) event; the reduced snapshot is what spec-kitty agent tasks status and merge preflight read.
The verdict field is validated against the canonical bridge specify_cli.status.event_verdicts() plus the proof-event extras commented, rejected, unknown (see specify_cli.proof.events). Do not invent verdict strings; an unknown verdict is rejected at the seam. Moving a WP to done requires a review triple with reviewer, verdict, and reference.
these CLI verbs — never by editing .md renders or the event log directly.
--review-feedback-file); an emptyrejection is not a review.
spk-run-review-wp performs the review that produces the verdict; this skill is only thecapture step. spk-run-implement-review drives the loop across WPs.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 11,086 | 7,768 | -30% | 1 | 1 | 0% | 889 | 958 | +8% | 0 | 0 | — |
case-02 | fail→pass | 7,311 | 2,789 | -62% | 1 | 1 | 0% | 272 | 1,007 | +270% | 0 | 0 | — |
case-03 | fail→fail | 23,775 | 10,166 | -57% | 1 | 1 | 0% | 702 | 1,099 | +57% | 0 | 0 | — |
case-04 | fail→pass | 13,576 | 28,109 | +107% | 1 | 1 | 0% | 1,163 | 2,838 | +144% | 0 | 0 | — |
case-05 | fail→fail | 16,026 | 21,217 | +32% | 1 | 1 | 0% | 215 | 922 | +329% | 0 | 0 | — |
case-06 | pass→pass | 8,221 | 7,693 | -6% | 1 | 1 | 0% | 1,060 | 1,609 | +52% | 0 | 0 | — |
case-07 | pass→pass | 8,275 | 6,993 | -15% | 1 | 1 | 0% | 1,126 | 1,268 | +13% | 0 | 0 | — |
case-08 | pass→pass | 10,942 | 4,650 | -58% | 1 | 1 | 0% | 1,654 | 1,158 | -30% | 0 | 0 | — |
case-09 | fail→pass | 9,188 | 4,784 | -48% | 1 | 1 | 0% | 1,275 | 1,278 | +0% | 0 | 0 | — |
case-10 | fail→pass | 31,667 | 8,631 | -73% | 1 | 1 | 0% | 1,966 | 1,838 | -7% | 0 | 0 | — |
case-11 | fail→pass | 19,304 | 3,050 | -84% | 1 | 1 | 0% | 2,148 | 1,035 | -52% | 0 | 0 | — |
case-12 | fail→pass | 10,329 | 5,080 | -51% | 1 | 1 | 0% | 1,665 | 1,466 | -12% | 0 | 0 | — |
case-13 | fail→pass | 11,398 | 5,102 | -55% | 1 | 1 | 0% | 1,674 | 1,128 | -33% | 0 | 0 | — |
case-14 | fail→pass | 12,678 | 58,664 | +363% | 1 | 1 | 0% | 2,124 | 1,065 | -50% | 0 | 0 | — |
case-15 | fail→pass | 17,422 | 3,413 | -80% | 1 | 1 | 0% | 2,791 | 1,115 | -60% | 0 | 0 | — |
case-16 | fail→pass | 10,282 | 6,199 | -40% | 1 | 1 | 0% | 1,715 | 1,525 | -11% | 0 | 0 | — |
case-17 | pass→pass | 11,121 | 16,621 | +49% | 1 | 1 | 0% | 1,476 | 1,474 | -0% | 0 | 0 | — |
case-18 | pass→pass | 12,048 | 3,731 | -69% | 1 | 1 | 0% | 1,772 | 1,136 | -36% | 0 | 0 | — |
case-19 | fail→pass | 20,864 | 8,393 | -60% | 1 | 1 | 0% | 3,176 | 1,209 | -62% | 0 | 0 | — |
case-20 | pass→pass | 8,676 | 35,852 | +313% | 1 | 1 | 0% | 887 | 1,219 | +37% | 0 | 0 | — |
case-21 | fail→pass | 35,230 | 2,349 | -93% | 1 | 1 | 0% | 1,703 | 947 | -44% | 0 | 0 | — |
case-22 | fail→pass | 10,759 | 3,596 | -67% | 1 | 1 | 0% | 1,632 | 1,065 | -35% | 0 | 0 | — |
case-23 | pass→pass | 46,064 | 21,419 | -54% | 1 | 1 | 0% | 2,051 | 2,368 | +15% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 20 counted toward the lift figure. The other 3 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +61 percentage points is the difference between those two pass rates over the 20 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.