Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when about to claim work is complete, fixed, or passing, before committing, before reporting a task done, or before telling the evaluator the batch is ready. Requires running the verification command and reading its output in this turn before any success claim; evidence before assertions always.
.claude/skills/fradser-verification-before-completion/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-14 | ✗→✓ | ▲ Improved | 11% | 0% |
| case-20 | ✓→✗ | ▼ Worse | -36% | 0% |
| case-22 | ✓→✗ | ▼ Worse | -41% | 0% |
| case-21 | ✓→✓ | = Same ✓ | 51% | 0% |
| case-01 | ✗→✗ | = Same ✗ | 486% | 0% |
NO COMPLETION CLAIMS WITHOUT FRESH VERIFICATION EVIDENCEIf you have not run the verification command in this turn, you cannot claim it passes. The superpowers-evaluator reads narration and runs commands independently — your unverified "done" can be waved through unless it self-evidences.
Violating the letter of this rule is violating the spirit of this rule.
Run this before claiming any status, expressing satisfaction, or returning a PASS-shaped verdict to the coordinator:
## Verification Commands)Skipping any step is lying, not verifying.
| Claim | Requires | Not Sufficient | |-------|----------|----------------| | Tests pass | Test command output: 0 failures | Previous run, "should pass" | | Linter clean | Linter output: 0 errors | Partial check, extrapolation | | Build succeeds | Build command: exit 0 | Linter passing, logs look good | | Bug fixed | Original-symptom test passes | Code changed, assumed fixed | | Regression test works | Red-green cycle verified | Test passes once | | Sub-agent completed | VCS diff shows changes + your re-run | Sub-agent reports "success" | | Requirements met | Line-by-line checklist vs spec | Tests passing alone |
| Excuse | Reality | |--------|---------| | "Should work now" | RUN the verification | | "I'm confident" | Confidence is not evidence | | "Just this once" | No exceptions | | "Linter passed" | Linter is not the test runner | | "The sub-agent said success" | Verify independently — sub-agents hallucinate done | | "Partial check is enough" | Partial proves nothing | | "Different words so the rule doesn't apply" | Spirit over letter | | "The evaluator will catch it anyway" | The evaluator reads narration; an unverified "done" can slip past. Self-verify first. |
This skill is the implementer-side pre-gate. The superpowers-evaluator is the independent read-only post-gate that re-runs verification commands itself. They do not overlap:
An unverified "done" from you can be waved through by the evaluator (it reads narration, not always files). Self-evidence closes that gap. Do not lean on the evaluator as a safety net — produce the evidence yourself first.
Run the command. Read the output. Paste the evidence. THEN claim the result.
This is non-negotiable.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 5,008 | 9,307 | +86% | 1 | 1 | 0% | 182 | 1,067 | +486% | 0 | 0 | — |
case-16 | fail→fail | 5,565 | 4,892 | -12% | 1 | 1 | 0% | 745 | 1,203 | +61% | 0 | 0 | — |
case-02 | fail→fail | 9,013 | 7,667 | -15% | 1 | 1 | 0% | 1,297 | 2,102 | +62% | 0 | 0 | — |
case-03 | fail→fail | 8,084 | 7,379 | -9% | 1 | 1 | 0% | 1,185 | 1,910 | +61% | 0 | 0 | — |
case-04 | fail→fail | 5,623 | 5,580 | -1% | 1 | 1 | 0% | 905 | 1,673 | +85% | 0 | 0 | — |
case-05 | fail→fail | 3,529 | 4,858 | +38% | 1 | 1 | 0% | 514 | 1,094 | +113% | 0 | 0 | — |
case-06 | fail→fail | 9,667 | 5,459 | -44% | 1 | 1 | 0% | 1,345 | 1,659 | +23% | 0 | 0 | — |
case-07 | fail→fail | 5,593 | 4,605 | -18% | 1 | 1 | 0% | 270 | 1,069 | +296% | 0 | 0 | — |
case-08 | fail→fail | 8,215 | 18,167 | +121% | 1 | 1 | 0% | 1,173 | 3,700 | +215% | 0 | 0 | — |
case-09 | fail→fail | 7,691 | 5,149 | -33% | 1 | 1 | 0% | 1,250 | 1,595 | +28% | 0 | 0 | — |
case-10 | fail→fail | 6,275 | 7,209 | +15% | 1 | 1 | 0% | 1,031 | 1,997 | +94% | 0 | 0 | — |
case-11 | fail→fail | 2,794 | 12,940 | +363% | 1 | 1 | 0% | 397 | 2,620 | +560% | 0 | 0 | — |
case-12 | fail→fail | 5,812 | 3,458 | -41% | 1 | 1 | 0% | 920 | 1,347 | +46% | 0 | 0 | — |
case-13 | fail→fail | 5,476 | 5,779 | +6% | 1 | 1 | 0% | 765 | 1,120 | +46% | 0 | 0 | — |
case-14 | fail→pass | 9,423 | 5,170 | -45% | 1 | 1 | 0% | 1,431 | 1,589 | +11% | 0 | 0 | — |
case-15 | fail→fail | 12,697 | 10,240 | -19% | 1 | 1 | 0% | 406 | 1,948 | +380% | 0 | 0 | — |
case-17 | fail→fail | 7,985 | 4,922 | -38% | 1 | 1 | 0% | 1,310 | 1,118 | -15% | 0 | 0 | — |
case-18 | fail→fail | 5,419 | 5,198 | -4% | 1 | 1 | 0% | 788 | 1,057 | +34% | 0 | 0 | — |
case-19 | fail→fail | 3,105 | 4,192 | +35% | 1 | 1 | 0% | 445 | 1,087 | +144% | 0 | 0 | — |
case-20 | pass→fail | 8,744 | 5,075 | -42% | 1 | 1 | 0% | 1,736 | 1,117 | -36% | 0 | 0 | — |
case-21 | pass→pass | 19,649 | 29,547 | +50% | 1 | 1 | 0% | 3,679 | 5,567 | +51% | 0 | 0 | — |
case-22 | pass→fail | 11,295 | 6,129 | -46% | 1 | 1 | 0% | 2,167 | 1,282 | -41% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 10 counted toward the lift figure. The other 12 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of -5 percentage points is the difference between those two pass rates over the 10 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.