Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Batch-validate ALL findings through the 7-Question Gate. Kills weak findings in bulk. Usage: /triage
.claude/skills/h-mmer-triage/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-06 | ✗→✓ | ▲ Improved | -28% | 0% |
| case-04 | ✗→✓ | ▲ Improved | -19% | 0% |
| case-05 | ✗→✓ | ▲ Improved | -8% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 22% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 7% | 0% |
Batch triage all findings.
ALL validator agents dispatched by this command MUST use .
validator agent with the finding detailsTRIAGE RESULTS
═══════════════
# Finding Decision Reason
1 GraphQL schema leakage KILL Q7 Never-submit: introspection alone
2 Config exposure SayTech KILL Q7 SPA client config is by design
3 Internal service URLs KILL Q6 Not exploitable externally
4 IDOR on /api/users/{id} PASS Confirmed with real data
5 XSS on comments PASS Cookie theft PoC works
PASSED: 2 findings → ready for /report
KILLED: 3 findings → removed from queueuv run python3 ../../tools/brain.py record <target> exhausted "<finding>" "<kill reason>"/report or /validate for full PoC + evidenceBatch triage should reduce the queue aggressively.
For each finding, produce:
Do not average weak findings into a stronger story. Chain them only when one finding provides a capability the next finding consumes.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-06 | fail→pass | 10,865 | 5,156 | -53% | 1 | 1 | 0% | 1,859 | 1,336 | -28% | 0 | 0 | — |
case-01 | fail→fail | 3,828 | 4,188 | +9% | 1 | 1 | 0% | 253 | 648 | +156% | 0 | 0 | — |
case-02 | fail→fail | 4,095 | 3,902 | -5% | 1 | 1 | 0% | 171 | 595 | +248% | 0 | 0 | — |
case-03 | fail→fail | 15,213 | 14,605 | -4% | 1 | 1 | 0% | 1,462 | 1,478 | +1% | 0 | 0 | — |
case-04 | fail→pass | 10,372 | 5,384 | -48% | 1 | 1 | 0% | 1,714 | 1,380 | -19% | 0 | 0 | — |
case-05 | fail→pass | 9,987 | 5,385 | -46% | 1 | 1 | 0% | 1,390 | 1,277 | -8% | 0 | 0 | — |
case-07 | fail→pass | 5,399 | 5,368 | -1% | 1 | 1 | 0% | 949 | 1,159 | +22% | 0 | 0 | — |
case-08 | fail→pass | 12,257 | 4,386 | -64% | 1 | 1 | 0% | 1,082 | 1,158 | +7% | 0 | 0 | — |
case-09 | fail→pass | 13,448 | 6,728 | -50% | 1 | 1 | 0% | 1,731 | 988 | -43% | 0 | 0 | — |
case-10 | pass→pass | 12,729 | 4,420 | -65% | 1 | 1 | 0% | 2,002 | 1,152 | -42% | 0 | 0 | — |
case-11 | pass→pass | 8,475 | 3,980 | -53% | 1 | 1 | 0% | 1,558 | 856 | -45% | 0 | 0 | — |
case-12 | fail→pass | 14,166 | 2,396 | -83% | 1 | 1 | 0% | 2,501 | 839 | -66% | 0 | 0 | — |
case-13 | fail→pass | 14,823 | 4,213 | -72% | 1 | 1 | 0% | 2,409 | 860 | -64% | 0 | 0 | — |
case-14 | pass→pass | 9,202 | 9,241 | +0% | 1 | 1 | 0% | 1,517 | 2,006 | +32% | 0 | 0 | — |
case-15 | fail→pass | 5,620 | 2,606 | -54% | 1 | 1 | 0% | 966 | 901 | -7% | 0 | 0 | — |
case-16 | fail→pass | 19,281 | 16,058 | -17% | 1 | 1 | 0% | 1,582 | 1,635 | +3% | 0 | 0 | — |
case-17 | pass→pass | 7,898 | 4,872 | -38% | 1 | 1 | 0% | 1,260 | 615 | -51% | 0 | 0 | — |
case-18 | pass→pass | 9,971 | 7,809 | -22% | 1 | 1 | 0% | 1,640 | 951 | -42% | 0 | 0 | — |
case-19 | fail→pass | 9,705 | 3,353 | -65% | 1 | 1 | 0% | 1,615 | 942 | -42% | 0 | 0 | — |
case-20 | pass→pass | 8,201 | 6,902 | -16% | 1 | 1 | 0% | 1,984 | 2,028 | +2% | 0 | 0 | — |
case-21 | pass→pass | 7,535 | 6,507 | -14% | 1 | 1 | 0% | 1,830 | 1,822 | -0% | 0 | 0 | — |
case-22 | fail→pass | 9,792 | 8,473 | -13% | 1 | 1 | 0% | 1,898 | 1,927 | +2% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 20 counted toward the lift figure. The other 2 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +55 percentage points is the difference between those two pass rates over the 20 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.