Install any skill in seconds. Free to start, no credit card required.
Get Started Free →A comprehensive verification system for Kiro sessions.
.claude/skills/affaan-m-verification-loop/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | -44% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 2% | 0% |
| case-15 | ✗→✓ | ▲ Improved | 26% | 0% |
| case-16 | ✗→✓ | ▲ Improved | -29% | 0% |
| case-18 | ✗→✓ | ▲ Improved | 104% | 0% |
Claude Code 工作階段的完整驗證系統。
在以下情況呼叫此技能:
bash# 檢查專案是否建置 npm run build 2>&1 | tail -20 # 或 pnpm build 2>&1 | tail -20
如果建置失敗,停止並在繼續前修復。
bash# TypeScript 專案 npx tsc --noEmit 2>&1 | head -30 # Python 專案 pyright . 2>&1 | head -30
報告所有型別錯誤。繼續前修復關鍵錯誤。
bash# JavaScript/TypeScript npm run lint 2>&1 | head -30 # Python ruff check . 2>&1 | head -30
bash# 執行帶覆蓋率的測試 npm run test -- --coverage 2>&1 | tail -50 # 檢查覆蓋率門檻 # 目標:最低 80%
報告:
bash# 檢查密鑰 grep -rn "sk-" --include="*.ts" --include="*.js" . 2>/dev/null | head -10 grep -rn "api_key" --include="*.ts" --include="*.js" . 2>/dev/null | head -10 # 檢查 console.log grep -rn "console.log" --include="*.ts" --include="*.tsx" src/ 2>/dev/null | head -10
bash# 顯示變更內容 git diff --stat git diff HEAD~1 --name-only
審查每個變更的檔案:
執行所有階段後,產生驗證報告:
驗證報告
==================
建置: [PASS/FAIL]
型別: [PASS/FAIL](X 個錯誤)
Lint: [PASS/FAIL](X 個警告)
測試: [PASS/FAIL](X/Y 通過,Z% 覆蓋率)
安全性: [PASS/FAIL](X 個問題)
差異: [X 個檔案變更]
整體: [READY/NOT READY] for PR
待修復問題:
1. ...
2. ...對於長時間工作階段,每 15 分鐘或重大變更後執行驗證:
markdown設定心理檢查點: - 完成每個函式後 - 完成元件後 - 移至下一個任務前 執行:/verify
此技能補充 PostToolUse hooks 但提供更深入的驗證。 Hooks 立即捕捉問題;此技能提供全面審查。
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 18,784 | 9,439 | -50% | 1 | 1 | 0% | 3,434 | 1,933 | -44% | 0 | 0 | — |
case-02 | fail→fail | 15,554 | 2,458 | -84% | 1 | 1 | 0% | 2,756 | 1,153 | -58% | 0 | 0 | — |
case-03 | fail→fail | 13,154 | 4,380 | -67% | 1 | 1 | 0% | 2,528 | 1,454 | -42% | 0 | 0 | — |
case-04 | fail→fail | 11,946 | 8,566 | -28% | 1 | 1 | 0% | 1,992 | 2,337 | +17% | 0 | 0 | — |
case-05 | pass→pass | 5,584 | 2,883 | -48% | 1 | 1 | 0% | 1,073 | 1,326 | +24% | 0 | 0 | — |
case-06 | pass→pass | 10,566 | 7,926 | -25% | 1 | 1 | 0% | 1,921 | 2,229 | +16% | 0 | 0 | — |
case-07 | fail→pass | 12,985 | 8,010 | -38% | 1 | 1 | 0% | 2,148 | 2,199 | +2% | 0 | 0 | — |
case-08 | pass→pass | 11,709 | 1,673 | -86% | 1 | 1 | 0% | 1,993 | 1,052 | -47% | 0 | 0 | — |
case-21 | pass→pass | 11,508 | 10,469 | -9% | 1 | 1 | 0% | 2,499 | 2,904 | +16% | 0 | 0 | — |
case-09 | pass→pass | 6,002 | 6,129 | +2% | 1 | 1 | 0% | 1,159 | 1,946 | +68% | 0 | 0 | — |
case-10 | pass→pass | 20,741 | 20,664 | -0% | 1 | 1 | 0% | 2,103 | 3,108 | +48% | 0 | 0 | — |
case-11 | pass→pass | 6,549 | 3,334 | -49% | 1 | 1 | 0% | 1,218 | 1,418 | +16% | 0 | 0 | — |
case-12 | fail→fail | 5,621 | 3,822 | -32% | 1 | 1 | 0% | 1,068 | 1,469 | +38% | 0 | 0 | — |
case-13 | pass→pass | 13,835 | 4,555 | -67% | 1 | 1 | 0% | 1,426 | 1,670 | +17% | 0 | 0 | — |
case-14 | pass→pass | 7,026 | 3,353 | -52% | 1 | 1 | 0% | 1,148 | 1,384 | +21% | 0 | 0 | — |
case-15 | fail→pass | 12,393 | 10,798 | -13% | 1 | 1 | 0% | 2,029 | 2,553 | +26% | 0 | 0 | — |
case-16 | fail→pass | 7,797 | 1,533 | -80% | 1 | 1 | 0% | 1,444 | 1,021 | -29% | 0 | 0 | — |
case-17 | pass→pass | 10,217 | 4,644 | -55% | 1 | 1 | 0% | 1,746 | 1,634 | -6% | 0 | 0 | — |
case-18 | fail→pass | 3,931 | 2,685 | -32% | 1 | 1 | 0% | 626 | 1,279 | +104% | 0 | 0 | — |
case-19 | fail→pass | 6,304 | 3,551 | -44% | 1 | 1 | 0% | 1,102 | 1,350 | +23% | 0 | 0 | — |
case-20 | pass→pass | 7,329 | 6,307 | -14% | 1 | 1 | 0% | 1,374 | 2,119 | +54% | 0 | 0 | — |
case-22 | pass→pass | 12,209 | 9,400 | -23% | 1 | 1 | 0% | 2,898 | 2,893 | -0% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +27 percentage points is the difference between those two pass rates over the 22 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.