Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Validate a finding through the 7-Question Gate + 4 gates. Kills weak findings FAST. Usage: /validate <finding description>
.claude/skills/h-mmer-validate/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-10 | ✗→✓ | ▲ Improved | 5% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -50% | 0% |
| case-08 | ✗→✓ | ▲ Improved | -39% | 0% |
| case-09 | ✗→✓ | ▲ Improved | -50% | 0% |
| case-11 | ✗→✓ | ▲ Improved | -41% | 0% |
Validate finding: $ARGUMENTS
This is the MOST IMPORTANT command. Run it BEFORE writing any report. It takes 30 seconds to kill a bad lead. A report takes 30 minutes.
Read findings.md and brain data. Locate the finding matching "$ARGUMENTS". Show finding details and ask user to confirm.
Launch validator agent: "Validate this finding through the 7-Question Gate and 4-gate checklist: finding details]. Check rules/hunting.md Rule 19 for the never-submit list AND rules/mistakes.md (REPORTING + METHODOLOGY sections) for lessons agents commonly miss — especially: (a) theoretical vs confirmed exploits, (b) file-path hallucinations, (c) CVSS-version mismatch per platform, (d) status-code asymmetry ≠ proven bug, (e) single-account IDOR ≠ cross-account leak. Output PASS, KILL, DOWNGRADE, or CHAIN REQUIRED with specific reason."
If PASS:
poc-builder agent to create minimal PoCuv run python3 ../../tools/capture.py screenshotreport-writer agent for platform-ready draftquality-check agent — block if score < 7/dupcheck <finding> then /submit <finding>If KILL:
uv run python3 ../../tools/brain.py record <target> exhausted "<finding>" "<kill reason>"/hunt <target> or /surface <target>If DOWNGRADE:
If CHAIN REQUIRED:
/chain to build the chainValidation is where mediocre hunters become expensive or elite.
Apply these hard checks before PASS:
If one check fails, prefer KILL or DOWNGRADE over "probably valid." Record the missing proof so the hunter can run one precise follow-up instead of re-litigating the whole bug.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 7,621 | 4,539 | -40% | 1 | 1 | 0% | 1,278 | 1,219 | -5% | 0 | 0 | — |
case-10 | fail→pass | 9,413 | 6,060 | -36% | 1 | 1 | 0% | 1,537 | 1,620 | +5% | 0 | 0 | — |
case-02 | fail→fail | 19,493 | 25,288 | +30% | 1 | 1 | 0% | 1,567 | 3,215 | +105% | 0 | 0 | — |
case-03 | fail→fail | 10,937 | 8,414 | -23% | 1 | 1 | 0% | 1,024 | 1,278 | +25% | 0 | 0 | — |
case-04 | fail→fail | 7,805 | 1,750 | -78% | 1 | 1 | 0% | 1,257 | 934 | -26% | 0 | 0 | — |
case-05 | fail→fail | 9,629 | 2,332 | -76% | 1 | 1 | 0% | 1,668 | 1,024 | -39% | 0 | 0 | — |
case-06 | pass→pass | 7,392 | 1,717 | -77% | 1 | 1 | 0% | 1,098 | 880 | -20% | 0 | 0 | — |
case-07 | fail→pass | 10,885 | 1,797 | -83% | 1 | 1 | 0% | 1,898 | 941 | -50% | 0 | 0 | — |
case-08 | fail→pass | 11,040 | 3,372 | -69% | 1 | 1 | 0% | 1,961 | 1,189 | -39% | 0 | 0 | — |
case-09 | fail→pass | 13,352 | 3,818 | -71% | 1 | 1 | 0% | 2,511 | 1,260 | -50% | 0 | 0 | — |
case-11 | fail→pass | 11,897 | 2,753 | -77% | 1 | 1 | 0% | 1,792 | 1,055 | -41% | 0 | 0 | — |
case-12 | fail→pass | 11,595 | 2,228 | -81% | 1 | 1 | 0% | 1,854 | 1,058 | -43% | 0 | 0 | — |
case-13 | fail→pass | 16,164 | 2,430 | -85% | 1 | 1 | 0% | 2,444 | 1,023 | -58% | 0 | 0 | — |
case-14 | pass→pass | 10,779 | 8,682 | -19% | 1 | 1 | 0% | 1,798 | 2,099 | +17% | 0 | 0 | — |
case-15 | pass→pass | 8,477 | 13,583 | +60% | 1 | 1 | 0% | 1,504 | 2,056 | +37% | 0 | 0 | — |
case-16 | fail→pass | 10,532 | 5,299 | -50% | 1 | 1 | 0% | 1,642 | 1,510 | -8% | 0 | 0 | — |
case-17 | pass→pass | 8,405 | 3,727 | -56% | 1 | 1 | 0% | 1,448 | 1,251 | -14% | 0 | 0 | — |
case-18 | pass→pass | 12,487 | 6,059 | -51% | 1 | 1 | 0% | 2,028 | 1,557 | -23% | 0 | 0 | — |
case-19 | fail→fail | 11,507 | 3,444 | -70% | 1 | 1 | 0% | 1,754 | 1,118 | -36% | 0 | 0 | — |
case-20 | pass→pass | 6,800 | 13,196 | +94% | 1 | 1 | 0% | 1,157 | 2,768 | +139% | 0 | 0 | — |
case-21 | fail→fail | 6,929 | 7,047 | +2% | 1 | 1 | 0% | 474 | 1,214 | +156% | 0 | 0 | — |
case-22 | pass→pass | 9,877 | 8,190 | -17% | 1 | 1 | 0% | 1,694 | 2,099 | +24% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +36 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.