Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Assess whether an agent can verify a small change without guessing or running an unnecessarily heavy loop
.claude/skills/kunanonj-cursor-plugin-agent-compat-agent-validation-review/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 19% | 0% |
| case-03 | ✗→✓ | ▲ Improved | -15% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 56% | 0% |
| case-05 | ✗→✓ | ▲ Improved | -53% | 0% |
| case-06 | ✗→✓ | ▲ Improved | -21% | 0% |
Checks whether an agent can verify a small change without falling back to a full-repo loop.
Use when the user wants to know whether an agent can safely verify its own work in a repo.
93/100 if there is a repeatable validation path and it gives useful signal, even if it is broader than ideal.84/100 if validation works but is heavier than it should be, repo-wide, or split across a few commands.68/100 if a valid loop probably exists but picking the right one takes guesswork or the output is too noisy to trust quickly.27/100 if there is no practical validation loop you can actually use.12/100 if the loop is blocked on secrets, accounts, or infrastructure you cannot reasonably access.83, 86, or 91 over a multiple of ten when that is the more honest read.Reply in plain text only (no markdown fences, no # headings, no emphasis syntax). Use this layout:
First line: Validation Loop Score: <score>/100
Then a short summary paragraph.
Then the line Problems followed by one bullet per line using - .
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | pass→fail | 7,727 | 3,527 | -54% | 1 | 1 | 0% | 1,255 | 816 | -35% | 0 | 0 | — |
case-02 | fail→pass | 7,675 | 5,638 | -27% | 1 | 1 | 0% | 1,382 | 1,644 | +19% | 0 | 0 | — |
case-03 | fail→pass | 10,673 | 4,213 | -61% | 1 | 1 | 0% | 1,757 | 1,485 | -15% | 0 | 0 | — |
case-04 | fail→pass | 4,922 | 3,959 | -20% | 1 | 1 | 0% | 858 | 1,341 | +56% | 0 | 0 | — |
case-05 | fail→pass | 14,744 | 3,660 | -75% | 1 | 1 | 0% | 2,664 | 1,263 | -53% | 0 | 0 | — |
case-06 | fail→pass | 9,529 | 4,562 | -52% | 1 | 1 | 0% | 1,780 | 1,414 | -21% | 0 | 0 | — |
case-07 | fail→pass | 13,660 | 5,797 | -58% | 1 | 1 | 0% | 2,378 | 1,666 | -30% | 0 | 0 | — |
case-08 | pass→pass | 8,618 | 7,121 | -17% | 1 | 1 | 0% | 1,500 | 1,815 | +21% | 0 | 0 | — |
case-09 | fail→pass | 11,767 | 3,884 | -67% | 1 | 1 | 0% | 1,956 | 1,362 | -30% | 0 | 0 | — |
case-10 | fail→pass | 14,908 | 5,988 | -60% | 1 | 1 | 0% | 2,495 | 1,705 | -32% | 0 | 0 | — |
case-11 | fail→pass | 10,410 | 6,700 | -36% | 1 | 1 | 0% | 1,810 | 1,686 | -7% | 0 | 0 | — |
case-12 | pass→pass | 9,515 | 5,797 | -39% | 1 | 1 | 0% | 1,601 | 1,595 | -0% | 0 | 0 | — |
case-13 | pass→pass | 11,682 | 6,143 | -47% | 1 | 1 | 0% | 1,906 | 1,699 | -11% | 0 | 0 | — |
case-14 | pass→pass | 11,787 | 4,463 | -62% | 1 | 1 | 0% | 2,237 | 1,443 | -35% | 0 | 0 | — |
case-15 | pass→pass | 9,381 | 5,921 | -37% | 1 | 1 | 0% | 1,792 | 1,822 | +2% | 0 | 0 | — |
case-16 | pass→pass | 10,488 | 7,869 | -25% | 1 | 1 | 0% | 1,856 | 2,030 | +9% | 0 | 0 | — |
case-17 | fail→pass | 10,408 | 6,803 | -35% | 1 | 1 | 0% | 1,923 | 1,742 | -9% | 0 | 0 | — |
case-18 | pass→pass | 10,783 | 5,769 | -46% | 1 | 1 | 0% | 2,088 | 1,720 | -18% | 0 | 0 | — |
case-19 | pass→pass | 7,759 | 4,866 | -37% | 1 | 1 | 0% | 1,547 | 1,529 | -1% | 0 | 0 | — |
case-20 | pass→pass | 18,140 | 18,221 | +0% | 1 | 1 | 0% | 3,893 | 4,475 | +15% | 0 | 0 | — |
case-21 | pass→pass | 9,030 | 8,058 | -11% | 1 | 1 | 0% | 2,463 | 2,642 | +7% | 0 | 0 | — |
case-22 | pass→pass | 12,934 | 20,931 | +62% | 1 | 1 | 0% | 3,107 | 5,396 | +74% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 21 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +41 percentage points is the difference between those two pass rates over the 21 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.