Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Verify completion and success claims with fresh evidence. Use before claiming a task is complete, a fix works, tests pass, or a feature is ready for GO.
.claude/skills/bilal140202-kiro-verify-completion/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 48% | 0% |
| case-02 | ✗→✓ | ▲ Improved | -16% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -7% | 0% |
| case-20 | ✗→✓ | ▲ Improved | -22% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 6% | 0% |
<background_information> This skill prevents false completion claims. A task, fix, or feature is only complete when supported by fresh evidence that matches the scope of the claim. </background_information>
<instructions>
GO from feature-level validationDo not use this skill for early planning or speculative status updates.
Provide:
TASKFIXTEST_OR_BUILDFEATURE_GOReturn one of:
VERIFIEDNOT_VERIFIEDMANUAL_VERIFY_REQUIREDAlso return:
Use the language specified in spec.json.
MANUAL_VERIFY_REQUIRED.Require:
Require:
Require:
Require:
A passing test suite alone is not enough for FEATURE_GO.
Return MANUAL_VERIFY_REQUIRED when:
Return NOT_VERIFIED when:
| Rationalization | Reality | |---|---| | “The subagent said it succeeded” | Reported success is not verification evidence. | | “Tests passed earlier” | Fresh evidence only. | | “Build should be fine because lint passed” | Lint does not prove build success. | | “Tests passed and build succeeded, so it must run” | Type erasure, module loading, native ABI, and boot-time config issues can still fail at runtime. | | “The feature is done because all tasks are checked off” | FEATURE_GO also requires coverage, integration, and design alignment. |
md## Verification Result - STATUS: VERIFIED | NOT_VERIFIED | MANUAL_VERIFY_REQUIRED - CLAIM_TYPE: TASK | FIX | TEST_OR_BUILD | FEATURE_GO - CLAIM: <exact claim> - EVIDENCE: <command/checklist and result> - GAPS: <scope/evidence mismatch or missing validation> - NOTES: <next action if not verified>
</instructions>
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 8,642 | 7,121 | -18% | 1 | 1 | 0% | 1,507 | 2,229 | +48% | 0 | 0 | — |
case-02 | fail→pass | 11,108 | 3,225 | -71% | 1 | 1 | 0% | 1,748 | 1,470 | -16% | 0 | 0 | — |
case-03 | fail→fail | 5,769 | 8,874 | +54% | 1 | 1 | 0% | 1,105 | 2,860 | +159% | 0 | 0 | — |
case-04 | pass→pass | 12,221 | 10,367 | -15% | 1 | 1 | 0% | 2,367 | 2,797 | +18% | 0 | 0 | — |
case-05 | pass→pass | 14,799 | 18,415 | +24% | 1 | 1 | 0% | 2,932 | 4,715 | +61% | 0 | 0 | — |
case-06 | pass→pass | 8,073 | 3,693 | -54% | 1 | 1 | 0% | 1,195 | 1,415 | +18% | 0 | 0 | — |
case-07 | fail→pass | 9,531 | 3,480 | -63% | 1 | 1 | 0% | 1,690 | 1,565 | -7% | 0 | 0 | — |
case-20 | fail→pass | 13,039 | 7,078 | -46% | 1 | 1 | 0% | 2,073 | 1,614 | -22% | 0 | 0 | — |
case-08 | fail→pass | 7,890 | 3,422 | -57% | 1 | 1 | 0% | 1,515 | 1,606 | +6% | 0 | 0 | — |
case-09 | fail→pass | 9,559 | 6,019 | -37% | 1 | 1 | 0% | 1,691 | 2,199 | +30% | 0 | 0 | — |
case-10 | fail→pass | 10,401 | 4,290 | -59% | 1 | 1 | 0% | 1,940 | 1,860 | -4% | 0 | 0 | — |
case-11 | fail→pass | 8,572 | 3,484 | -59% | 1 | 1 | 0% | 1,498 | 1,627 | +9% | 0 | 0 | — |
case-12 | fail→pass | 8,010 | 3,732 | -53% | 1 | 1 | 0% | 1,317 | 1,622 | +23% | 0 | 0 | — |
case-13 | fail→pass | 10,907 | 2,977 | -73% | 1 | 1 | 0% | 1,879 | 1,507 | -20% | 0 | 0 | — |
case-14 | fail→pass | 9,041 | 3,126 | -65% | 1 | 1 | 0% | 1,556 | 1,426 | -8% | 0 | 0 | — |
case-15 | fail→pass | 8,645 | 2,725 | -68% | 1 | 1 | 0% | 1,420 | 1,447 | +2% | 0 | 0 | — |
case-16 | fail→pass | 9,123 | 3,584 | -61% | 1 | 1 | 0% | 1,688 | 1,554 | -8% | 0 | 0 | — |
case-17 | fail→pass | 9,380 | 3,850 | -59% | 1 | 1 | 0% | 1,428 | 1,647 | +15% | 0 | 0 | — |
case-18 | fail→pass | 7,402 | 3,386 | -54% | 1 | 1 | 0% | 1,359 | 1,573 | +16% | 0 | 0 | — |
case-19 | fail→pass | 7,428 | 3,959 | -47% | 1 | 1 | 0% | 1,255 | 1,522 | +21% | 0 | 0 | — |
case-21 | fail→pass | 8,977 | 3,071 | -66% | 1 | 1 | 0% | 1,374 | 1,427 | +4% | 0 | 0 | — |
case-22 | fail→pass | 10,396 | 2,743 | -74% | 1 | 1 | 0% | 1,629 | 1,365 | -16% | 0 | 0 | — |
case-23 | fail→pass | 6,148 | 5,252 | -15% | 1 | 1 | 0% | 1,003 | 1,379 | +37% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of +83 percentage points is the difference between those two pass rates over the 23 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.