Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Review a task implementation against approved specs, task boundaries, and verification evidence. Use after an implementer finishes a task, after remediation, or before accepting a task as complete.
.claude/skills/bilal140202-kiro-review/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 1316% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 142% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 135% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 54% | 0% |
| case-17 | ✗→✓ | ▲ Improved | 90% | 0% |
<background_information> This skill performs task-local adversarial review. It verifies that the implementation is real, complete, bounded, aligned with approved requirements and design, and supported by mechanical verification evidence.
Boundary terminology continuity:
Boundary CandidatesBoundary Commitments_Boundary:_Boundary Violations</background_information>
<instructions>
READY_FOR_REVIEW[x]Do not use this skill to invent missing requirements or silently reinterpret the spec.
Provide:
tasks.mdrequirements.md, design.md, optionally tasks.md)_Boundary:_ scope constraints## Implementation Notes entries when applicableReturn one of:
APPROVEDREJECTEDAlso return:
Use the language specified in spec.json.
Run git diff to inspect the actual code changes. If the diff is large or ambiguous, read the changed files directly. Do not trust the implementer report as source of truth.
Read the spec yourself. Read the diff yourself. Verify mechanically where possible. Reject on concrete failures rather than interpretive optimism. The main review question is not just "does it work?" but "does it stay inside the approved responsibility boundary without hiding new coupling?"
Run these checks and use the result as primary signal.
TBD, TODO, FIXME, HACK, XXX._Boundary:_ scope.RED_PHASE_OUTPUT.requirements.md.design.md.Use:
Critical for broken functionality, invalid verification, data loss, security risk, or major scope violationImportant for required fixes before acceptanceSuggestion for non-blocking improvementsFYI for informational notesEscalate instead of papering over the issue when:
| Rationalization | Reality | |---|---| | “Tests pass, so approve” | Passing tests do not prove spec compliance or boundary respect. | | “The extra behavior is useful” | Extra behavior outside approved scope is still drift. | | “The implementer said RED was done” | RED must be evidenced, not asserted. | | “This gap is small enough to let through” | Real gaps must be rejected or escalated. |
md## Review Verdict - VERDICT: APPROVED | REJECTED - TASK: <task-id> - MECHANICAL_RESULTS: - Tests: PASS | FAIL (command and exit code) - TBD/TODO grep: CLEAN | <count> matches - Secrets grep: CLEAN | <count> matches - Static checks: PASS | FAIL | SPOT_CHECKED - Boundary: WITHIN | <files outside boundary> - Boundary audit: CLEAN | <spillover / hidden dependency findings> - RED phase: VERIFIED | MISSING | N/A - FINDINGS: 1. <specific finding with exact files/spec refs> - REMEDIATION: <mandatory if REJECTED> - SUMMARY: <one sentence>
</instructions>
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-14 | pass→pass | 7,689 | 4,092 | -47% | 1 | 1 | 0% | 1,292 | 2,322 | +80% | 0 | 0 | — |
case-01 | fail→fail | 5,365 | 5,067 | -6% | 1 | 1 | 0% | 276 | 1,837 | +566% | 0 | 0 | — |
case-02 | fail→pass | 5,201 | 15,575 | +199% | 1 | 1 | 0% | 319 | 4,516 | +1316% | 0 | 0 | — |
case-03 | fail→pass | 13,358 | 15,748 | +18% | 1 | 1 | 0% | 1,501 | 3,639 | +142% | 0 | 0 | — |
case-04 | pass→pass | 16,573 | 19,833 | +20% | 1 | 1 | 0% | 2,858 | 4,737 | +66% | 0 | 0 | — |
case-05 | pass→pass | 16,646 | 21,873 | +31% | 1 | 1 | 0% | 3,140 | 5,519 | +76% | 0 | 0 | — |
case-06 | pass→fail | 21,415 | 20,843 | -3% | 1 | 1 | 0% | 3,754 | 5,389 | +44% | 0 | 0 | — |
case-07 | fail→pass | 7,707 | 7,738 | +0% | 1 | 1 | 0% | 1,223 | 2,873 | +135% | 0 | 0 | — |
case-08 | fail→pass | 9,249 | 5,222 | -44% | 1 | 1 | 0% | 1,538 | 2,376 | +54% | 0 | 0 | — |
case-09 | pass→pass | 20,369 | 17,813 | -13% | 1 | 1 | 0% | 1,276 | 2,708 | +112% | 0 | 0 | — |
case-15 | pass→pass | 11,697 | 6,680 | -43% | 1 | 1 | 0% | 1,855 | 2,657 | +43% | 0 | 0 | — |
case-10 | pass→pass | 9,723 | 7,588 | -22% | 1 | 1 | 0% | 1,513 | 2,985 | +97% | 0 | 0 | — |
case-11 | pass→pass | 9,921 | 5,223 | -47% | 1 | 1 | 0% | 1,623 | 2,635 | +62% | 0 | 0 | — |
case-12 | pass→pass | 11,510 | 6,558 | -43% | 1 | 1 | 0% | 1,890 | 2,853 | +51% | 0 | 0 | — |
case-13 | pass→pass | 11,806 | 5,304 | -55% | 1 | 1 | 0% | 2,040 | 2,433 | +19% | 0 | 0 | — |
case-16 | pass→pass | 8,016 | 6,228 | -22% | 1 | 1 | 0% | 1,464 | 2,678 | +83% | 0 | 0 | — |
case-17 | fail→pass | 9,392 | 8,304 | -12% | 1 | 1 | 0% | 1,603 | 3,040 | +90% | 0 | 0 | — |
case-18 | pass→pass | 8,800 | 5,730 | -35% | 1 | 1 | 0% | 1,571 | 2,597 | +65% | 0 | 0 | — |
case-19 | fail→pass | 8,109 | 10,127 | +25% | 1 | 1 | 0% | 1,314 | 3,019 | +130% | 0 | 0 | — |
case-20 | fail→pass | 12,832 | 8,034 | -37% | 1 | 1 | 0% | 1,877 | 2,934 | +56% | 0 | 0 | — |
case-21 | pass→pass | 8,884 | 5,368 | -40% | 1 | 1 | 0% | 1,430 | 2,426 | +70% | 0 | 0 | — |
case-22 | fail→pass | 6,779 | 1,902 | -72% | 1 | 1 | 0% | 1,038 | 1,815 | +75% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 20 counted toward the lift figure. The other 2 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +29 percentage points is the difference between those two pass rates over the 20 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.