Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Audit and slim an existing Agent Skill by separating durable capability delta from generic model knowledge, stale facts, repeated warnings, and ritual. Use when creating, updating, reviewing, or modernizing a SKILL.md; reducing instruction or token burden; resolving overlapping triggers; or deciding whether to keep, trim, measure, or delete a skill. Do not use to shorten ordinary prose or source code unrelated to agent skills, run an upgrade-only compatibility canary, or prove an efficacy claim;
.claude/skills/paranoidandroid2124-skill-issue/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-06 | ✗→✓ | ▲ Improved | 270% | 0% |
| case-08 | ✗→✓ | ▲ Improved | -3% | 0% |
| case-09 | ✗→✓ | ▲ Improved | -19% | 0% |
| case-13 | ✗→✓ | ▲ Improved | -4% | 0% |
| case-16 | ✗→✓ | ▲ Improved | 35% | 0% |
Reduce the skill to the smallest intervention that demonstrably improves its target task.
Assume the model is innocent until the skill proves otherwise.
SKILL.md and every resource it directly routes to.observable completion condition.
node scripts/audit-skill.mjs <skill-directory> from this skilldirectory when Node.js is available. Treat its findings as leads, not semantic verdicts.
not infer this from reputation or model version.
Assign every material instruction one disposition:
boundaries, deterministic operations, output contracts, and observable verification.
model or tool assumptions, and process whose cost exceeds the failure risk.
current environment, or no measurable benefit over baseline.
Read references/decision-rubric.md when a classification is disputed or the skill mixes domain knowledge and workflow.
Never trim input validation at trust boundaries, security controls, accessibility basics, error handling that prevents data loss, or an explicit user requirement.
name and description. Preserve recognized optional orharness-specific fields when they carry licensing, compatibility, tool restrictions, security boundaries, or useful metadata. Remove a field only after classifying its behavior and portability cost.
description.references/; do notduplicate it in the body.
when the standard library suffices.
audit-only request, return the patch proposal without changing files.
audit-skill.mjs and report before/after lines, words, estimatedtokens, resources, and unresolved findings. Include the exact command and counting method; retain the raw report or its path when results will be compared.
$trust-me-bro when available.Otherwise run the same pinned task once without the skill and once with it, keep contexts isolated, and verify outputs from artifacts rather than self-reports.
Lead with keep, trim, delete, or measure. Include:
Stars, prose length, and the implementer's confidence are not evidence.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-18 | pass→pass | 11,215 | 10,858 | -3% | 1 | 1 | 0% | 1,870 | 2,316 | +24% | 0 | 0 | — |
case-01 | fail→fail | 5,999 | 6,310 | +5% | 1 | 1 | 0% | 471 | 1,219 | +159% | 0 | 0 | — |
case-02 | fail→fail | 6,265 | 6,516 | +4% | 1 | 1 | 0% | 409 | 1,063 | +160% | 0 | 0 | — |
case-03 | fail→fail | 17,792 | 5,256 | -70% | 1 | 1 | 0% | 3,193 | 957 | -70% | 0 | 0 | — |
case-04 | pass→pass | 15,150 | 17,537 | +16% | 1 | 1 | 0% | 3,114 | 3,997 | +28% | 0 | 0 | — |
case-05 | pass→pass | 13,828 | 10,689 | -23% | 1 | 1 | 0% | 2,321 | 2,446 | +5% | 0 | 0 | — |
case-06 | fail→pass | 5,217 | 16,082 | +208% | 1 | 1 | 0% | 868 | 3,210 | +270% | 0 | 0 | — |
case-07 | fail→fail | 24,069 | 6,254 | -74% | 1 | 1 | 0% | 6,394 | 1,092 | -83% | 0 | 0 | — |
case-08 | fail→pass | 12,014 | 6,147 | -49% | 1 | 1 | 0% | 1,825 | 1,778 | -3% | 0 | 0 | — |
case-09 | fail→pass | 8,458 | 3,452 | -59% | 1 | 1 | 0% | 1,669 | 1,357 | -19% | 0 | 0 | — |
case-10 | fail→fail | 8,053 | 3,285 | -59% | 1 | 1 | 0% | 1,324 | 1,227 | -7% | 0 | 0 | — |
case-11 | fail→fail | 13,709 | 6,175 | -55% | 1 | 1 | 0% | 2,523 | 1,783 | -29% | 0 | 0 | — |
case-12 | pass→pass | 10,287 | 5,623 | -45% | 1 | 1 | 0% | 1,577 | 1,698 | +8% | 0 | 0 | — |
case-13 | fail→pass | 10,231 | 5,553 | -46% | 1 | 1 | 0% | 1,813 | 1,743 | -4% | 0 | 0 | — |
case-14 | pass→pass | 11,931 | 5,421 | -55% | 1 | 1 | 0% | 1,829 | 1,647 | -10% | 0 | 0 | — |
case-15 | pass→pass | 9,111 | 9,762 | +7% | 1 | 1 | 0% | 1,404 | 2,518 | +79% | 0 | 0 | — |
case-16 | fail→pass | 6,532 | 3,138 | -52% | 1 | 1 | 0% | 979 | 1,319 | +35% | 0 | 0 | — |
case-17 | fail→pass | 7,403 | 3,735 | -50% | 1 | 1 | 0% | 1,224 | 1,444 | +18% | 0 | 0 | — |
case-19 | fail→fail | 4,041 | 3,100 | -23% | 1 | 1 | 0% | 553 | 1,095 | +98% | 0 | 0 | — |
case-20 | fail→pass | 10,293 | 6,384 | -38% | 1 | 1 | 0% | 1,583 | 1,754 | +11% | 0 | 0 | — |
case-21 | fail→pass | 17,336 | 8,823 | -49% | 1 | 1 | 0% | 3,021 | 2,222 | -26% | 0 | 0 | — |
case-22 | fail→pass | 9,247 | 5,981 | -35% | 1 | 1 | 0% | 1,537 | 1,877 | +22% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 18 counted toward the lift figure. The other 4 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +41 percentage points is the difference between those two pass rates over the 18 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.