Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Run a small pinned regression suite to detect whether an existing Agent Skill's triggers, tool routing, permissions, output contract, or versioned guidance drifted after a model, harness, tool, or API update. Use when upgrading a model or agent, checking a release, scheduling compatibility checks, or investigating why a formerly working skill changed behavior. Do not use for general skill slimming or to prove broad efficacy; those need a structural audit or paired trial.
.claude/skills/paranoidandroid2124-this-is-fine/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-07 | ✗→✓ | ▲ Improved | -18% | 0% |
| case-01 | ✗→✓ | ▲ Improved | 57% | 0% |
| case-02 | ✗→✓ | ▲ Improved | -35% | 0% |
| case-03 | ✗→✓ | ▲ Improved | -37% | 0% |
| case-08 | ✗→✓ | ▲ Improved | -38% | 0% |
Detect compatibility drift cheaply. Do not turn a canary into a full benchmark.
The coffee is optional. The pinned regression cases are not.
Otherwise pin the new environment and compare against stored expected observations.
references/canary-contract.md.
externally.
Keep task inputs, permissions, tools, reasoning, and timeout stable. Use fresh contexts and capture raw evidence for:
Run the last known-good environment first when practical, then the candidate environment. A missing old runtime is a limitation, not permission to invent a comparison.
sources.
changes.
the run green.
$trust-me-bro when available before claiming the new version improvesefficacy. Otherwise compare the same pinned task in isolated baseline and candidate contexts with one declared verifier. A canary alone never proves uplift.
Report the environment delta, case-by-case observation, raw evidence, drift class, severity, smallest proposed fix, and whether release should pass, warn, or block.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-07 | fail→pass | 12,864 | 7,328 | -43% | 1 | 1 | 0% | 2,161 | 1,776 | -18% | 0 | 0 | — |
case-01 | fail→pass | 15,915 | 21,584 | +36% | 1 | 1 | 0% | 3,008 | 4,718 | +57% | 0 | 0 | — |
case-02 | fail→pass | 28,629 | 15,644 | -45% | 1 | 1 | 0% | 5,251 | 3,424 | -35% | 0 | 0 | — |
case-03 | fail→pass | 25,465 | 11,842 | -53% | 1 | 1 | 0% | 4,460 | 2,794 | -37% | 0 | 0 | — |
case-04 | pass→fail | 20,828 | 14,404 | -31% | 1 | 1 | 0% | 3,149 | 2,999 | -5% | 0 | 0 | — |
case-05 | pass→pass | 24,801 | 21,391 | -14% | 1 | 1 | 0% | 3,869 | 4,135 | +7% | 0 | 0 | — |
case-06 | pass→pass | 16,086 | 8,379 | -48% | 1 | 1 | 0% | 2,621 | 2,058 | -21% | 0 | 0 | — |
case-08 | fail→pass | 12,129 | 3,572 | -71% | 1 | 1 | 0% | 1,850 | 1,156 | -38% | 0 | 0 | — |
case-09 | pass→pass | 14,101 | 6,729 | -52% | 1 | 1 | 0% | 2,218 | 1,502 | -32% | 0 | 0 | — |
case-10 | pass→pass | 10,550 | 5,735 | -46% | 1 | 1 | 0% | 1,746 | 1,370 | -22% | 0 | 0 | — |
case-11 | fail→pass | 9,327 | 4,761 | -49% | 1 | 1 | 0% | 1,558 | 1,304 | -16% | 0 | 0 | — |
case-12 | fail→pass | 15,198 | 4,589 | -70% | 1 | 1 | 0% | 2,401 | 1,363 | -43% | 0 | 0 | — |
case-13 | pass→pass | 9,466 | 2,682 | -72% | 1 | 1 | 0% | 1,233 | 962 | -22% | 0 | 0 | — |
case-14 | fail→pass | 6,756 | 2,471 | -63% | 1 | 1 | 0% | 1,121 | 920 | -18% | 0 | 0 | — |
case-15 | fail→pass | 7,158 | 2,515 | -65% | 1 | 1 | 0% | 1,058 | 897 | -15% | 0 | 0 | — |
case-16 | pass→pass | 10,007 | 3,222 | -68% | 1 | 1 | 0% | 1,646 | 1,041 | -37% | 0 | 0 | — |
case-17 | pass→pass | 8,932 | 4,358 | -51% | 1 | 1 | 0% | 1,424 | 1,196 | -16% | 0 | 0 | — |
case-18 | pass→pass | 11,175 | 5,924 | -47% | 1 | 1 | 0% | 1,689 | 1,425 | -16% | 0 | 0 | — |
case-19 | fail→pass | 11,312 | 2,251 | -80% | 1 | 1 | 0% | 1,809 | 849 | -53% | 0 | 0 | — |
case-20 | fail→pass | 8,961 | 4,829 | -46% | 1 | 1 | 0% | 1,529 | 1,182 | -23% | 0 | 0 | — |
case-21 | fail→pass | 7,799 | 3,874 | -50% | 1 | 1 | 0% | 1,297 | 933 | -28% | 0 | 0 | — |
case-22 | pass→pass | 9,358 | 5,391 | -42% | 1 | 1 | 0% | 1,585 | 1,246 | -21% | 0 | 0 | — |
case-23 | fail→pass | 19,349 | 10,949 | -43% | 1 | 1 | 0% | 2,767 | 2,098 | -24% | 0 | 0 | — |
case-24 | fail→pass | 11,208 | 1,810 | -84% | 1 | 1 | 0% | 1,857 | 796 | -57% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 24 cases were attempted. The headline lift of +54 percentage points is the difference between those two pass rates over the 24 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.