Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Review runtime-owned outputs using the Spec Kitty review workflow surface, then direct approval or rejection with structured feedback. Triggers: "review this work package", "check runtime output", "approve this step", "review WP", "is this WP ready to approve", "check this implementation". Does NOT handle: setup-only repair requests, direct implementation work, editorial glossary maintenance, or runtime loop advancement.
.claude/skills/priivacy-ai-spec-kitty-runtime-review/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-07 | ✗→✓ | ▲ Improved | -12% | 0% |
| case-08 | ✗→✓ | ▲ Improved | -28% | 0% |
| case-09 | ✗→✓ | ▲ Improved | -3% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 27% | 0% |
| case-22 | ✗→✓ | ▲ Improved | 42% | 0% |
Operate the Spec Kitty review workflow surface: load review context, claim a work package, read the generated review prompt, and issue an approve or reject transition.
bashspec-kitty charter context --action review --json
The returned text contains governance context. The review prompt (generated in Step 2) includes project-specific acceptance criteria and review guidance from charter.offering — do not restate those rules here.
bash# Claim a specific WP (or omit WP## to auto-select from for_review lane) spec-kitty agent action review WP## --agent <your-name>
This moves the WP from for_review to in_review and prints the path to a generated review prompt file. Read that path from the command output.
bashcat <prompt-file-path>
The review prompt contains:
main)Follow the review prompt. It is the source of truth for what to check and how to check it. The review criteria come from charter.offering and the WP definition, not from this skill.
If the mission ships a kitty-specs/<mission>/contracts/ directory, walk it before issuing a verdict. Contracts pin concrete examples (payloads, CLI invocations, schema fragments, error messages, allowed-value sets) that the implementation must round-trip verbatim. Vocabulary or shape drift here is invisible to spec-level review and surfaces late at mission-review or downstream consumption time.
ls kitty-specs/<mission>/contracts/. If the directory isabsent, record a one-line note ("no contracts/ artifact") and skip to Step 4.
symbols, payloads, or commands it pins. Skipped files get a one-line note ("orthogonal to WP##").
CLI invocation, schema fragment, and error message in the in-scope contracts. Verify the implementation accepts inputs and produces outputs verbatim: singular vs plural, default values, enum members, exit codes, flag names — all of it.
registered triggers, enum members), require one assertion that mirrors the runtime constant against its test-side or doc-side copy byte-for- byte. If the WP introduces such a constant without that mirror assertion, reject.
example is a BLOCKER, cited as contract file:line plus impl symbol. Contracts are spec artifacts, not preferences.
Take exactly one action — never "approve with conditions".
bashspec-kitty agent tasks move-task WP## --to approved --note "Review passed: <summary>"
Write structured feedback to a temp file, then move the WP back to planned:
bashspec-kitty agent tasks move-task WP## --to planned --force \ --review-feedback-file <feedback-file-path>
Every blocking finding must map to a specific, verifiable remediation action.
> The feedback reference is mandatory, not optional. A rejection moves the WP > along a backward "review-rejection" edge, and the status contract requires that > edge to carry a rationale: the --review-feedback-file (or --note) you > pass is recorded as the transition's review_ref / reason and travels on the > event wire. Rejecting a WP to planned/in_progress without a feedback > reference or note produces a contract-invalid status event that is accepted > locally but silently rejected by hosted sync — so the rejection never > propagates. Always attach the feedback file (or a --note) when you reject or > send a WP back for rework. The rework agent reads exactly that referenced > feedback to know what to fix.
If you rejected and the WP has downstream dependents:
bashspec-kitty agent tasks status --mission <mission-slug>
Note dependent WPs and include a rebase warning in your feedback.
passes even if the reviewer would have done it differently
checks, criteria, and doctrine context for this WP
the original implementing agent
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 14,400 | 15,085 | +5% | 1 | 1 | 0% | 283 | 1,519 | +437% | 0 | 0 | — |
case-02 | fail→fail | 26,152 | 14,500 | -45% | 1 | 1 | 0% | 3,596 | 1,550 | -57% | 0 | 0 | — |
case-03 | fail→fail | 24,708 | 18,294 | -26% | 1 | 1 | 0% | 1,548 | 1,500 | -3% | 0 | 0 | — |
case-04 | fail→fail | 20,712 | 60,957 | +194% | 1 | 1 | 0% | 2,895 | 5,359 | +85% | 0 | 0 | — |
case-05 | pass→fail | 14,400 | 10,693 | -26% | 1 | 1 | 0% | 1,717 | 3,052 | +78% | 0 | 0 | — |
case-06 | pass→fail | 21,686 | 15,148 | -30% | 1 | 1 | 0% | 2,940 | 1,500 | -49% | 0 | 0 | — |
case-07 | fail→pass | 14,211 | 6,468 | -54% | 1 | 1 | 0% | 1,643 | 1,439 | -12% | 0 | 0 | — |
case-08 | fail→pass | 17,223 | 6,945 | -60% | 1 | 1 | 0% | 2,167 | 1,565 | -28% | 0 | 0 | — |
case-09 | fail→pass | 16,367 | 8,281 | -49% | 1 | 1 | 0% | 1,844 | 1,784 | -3% | 0 | 0 | — |
case-10 | pass→pass | 15,303 | 7,758 | -49% | 1 | 1 | 0% | 1,787 | 1,764 | -1% | 0 | 0 | — |
case-11 | pass→pass | 14,477 | 8,898 | -39% | 1 | 1 | 0% | 1,533 | 1,903 | +24% | 0 | 0 | — |
case-12 | pass→pass | 10,127 | 9,270 | -8% | 1 | 1 | 0% | 811 | 2,116 | +161% | 0 | 0 | — |
case-13 | fail→pass | 15,001 | 10,453 | -30% | 1 | 1 | 0% | 1,687 | 2,144 | +27% | 0 | 0 | — |
case-22 | fail→pass | 11,801 | 7,517 | -36% | 1 | 1 | 0% | 1,181 | 1,679 | +42% | 0 | 0 | — |
case-14 | fail→pass | 11,908 | 7,065 | -41% | 1 | 1 | 0% | 1,139 | 1,565 | +37% | 0 | 0 | — |
case-15 | fail→pass | 33,245 | 8,235 | -75% | 1 | 1 | 0% | 2,033 | 1,822 | -10% | 0 | 0 | — |
case-16 | fail→pass | 10,976 | 9,068 | -17% | 1 | 1 | 0% | 1,027 | 1,971 | +92% | 0 | 0 | — |
case-17 | fail→pass | 13,085 | 9,761 | -25% | 1 | 1 | 0% | 1,205 | 2,083 | +73% | 0 | 0 | — |
case-18 | fail→pass | 16,491 | 6,896 | -58% | 1 | 1 | 0% | 1,840 | 1,553 | -16% | 0 | 0 | — |
case-19 | pass→pass | 15,623 | 9,074 | -42% | 1 | 1 | 0% | 1,613 | 1,993 | +24% | 0 | 0 | — |
case-20 | fail→pass | 15,335 | 8,932 | -42% | 1 | 1 | 0% | 1,509 | 1,941 | +29% | 0 | 0 | — |
case-21 | fail→pass | 16,530 | 8,215 | -50% | 1 | 1 | 0% | 1,872 | 1,796 | -4% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 17 counted toward the lift figure. The other 5 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +45 percentage points is the difference between those two pass rates over the 17 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
| Model | Method | Date | Lift |
|---|---|---|---|
| gemini-3.6-flash | verified | 8/17/2026 | +64% |
| gemini-3.6-flash | verified | 8/13/2026 | +38% |
Other measured skills in the registry, with their headline benchmark lift.