Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when the user wants to turn a script, repeated workflow, runbook, or hard-won procedure into a reusable agent skill, or asks "make this a skill", "extract this into a skill", or "I keep doing this manually".
.claude/skills/escoffier-labs-skillify/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | -13% | 0% |
| case-03 | ✗→✓ | ▲ Improved | -17% | 0% |
| case-01 | ✗→✓ | ▲ Improved | 21% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -2% | 0% |
| case-09 | ✗→✓ | ▲ Improved | -2% | 0% |
Extract a procedure you keep repeating into a SKILL.md an agent can discover and follow. The input is a script, a shell history pattern, a runbook, or "the thing we just did"; the output is an installable skill.
Make a skill when the technique was not obvious the first time, will recur across projects, and involves judgment. Do not make a skill for one-off fixes, things a linter or script can fully enforce (automate those instead), or project-specific conventions (those belong in the project's CLAUDE.md or AGENTS.md).
scripts/ next to SKILL.md and have the skill invoke it; keep judgment in the prose.markdown--- name: verb-first-hyphenated-name # or the tool's own name when wrapping a named tool description: Use when [triggering conditions and symptoms only, third person, under 500 chars] --- # name One-paragraph overview: what this achieves and the core principle. ## Steps / Pattern (the procedure) ## Rules (the constraints that protect against known failures) ## Common mistakes (what went wrong historically, and the fix)
The description field is load-bearing: agents read it to decide whether to load the skill. Describe WHEN to use it (symptoms, triggers, situations), never summarize the workflow, or agents will follow the one-line summary instead of reading the skill.
~/.claude/skills/ for Claude Code).claude/skills/ in the repoThe same SKILL.md format works across harnesses that support agent skills; only the install location differs.
Hand the skill to a fresh agent (subagent or new session) with a realistic task that should trigger it. If the agent misapplies a step or asks a question the skill should have answered, the skill has a gap; fix it and re-test. An untested skill is a guess.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-02 | fail→pass | 40,562 | 13,493 | -67% | 1 | 1 | 0% | 3,504 | 3,031 | -13% | 0 | 0 | — |
case-03 | fail→pass | 18,890 | 11,503 | -39% | 1 | 1 | 0% | 3,203 | 2,665 | -17% | 0 | 0 | — |
case-01 | fail→pass | 17,186 | 16,690 | -3% | 1 | 1 | 0% | 3,071 | 3,707 | +21% | 0 | 0 | — |
case-04 | pass→pass | 10,369 | 8,872 | -14% | 1 | 1 | 0% | 1,922 | 2,232 | +16% | 0 | 0 | — |
case-05 | pass→pass | 10,838 | 3,751 | -65% | 1 | 1 | 0% | 1,830 | 1,376 | -25% | 0 | 0 | — |
case-06 | pass→pass | 7,886 | 5,683 | -28% | 1 | 1 | 0% | 1,360 | 1,765 | +30% | 0 | 0 | — |
case-07 | fail→pass | 15,645 | 11,927 | -24% | 1 | 1 | 0% | 2,764 | 2,698 | -2% | 0 | 0 | — |
case-08 | pass→pass | 5,427 | 3,768 | -31% | 1 | 1 | 0% | 1,041 | 1,362 | +31% | 0 | 0 | — |
case-09 | fail→pass | 7,249 | 2,866 | -60% | 1 | 1 | 0% | 1,223 | 1,202 | -2% | 0 | 0 | — |
case-10 | fail→pass | 7,195 | 3,779 | -47% | 1 | 1 | 0% | 1,486 | 1,395 | -6% | 0 | 0 | — |
case-11 | fail→pass | 7,096 | 2,966 | -58% | 1 | 1 | 0% | 1,329 | 1,225 | -8% | 0 | 0 | — |
case-12 | pass→pass | 6,401 | 6,208 | -3% | 1 | 1 | 0% | 970 | 1,919 | +98% | 0 | 0 | — |
case-13 | pass→pass | 12,877 | 7,059 | -45% | 1 | 1 | 0% | 2,243 | 1,949 | -13% | 0 | 0 | — |
case-14 | pass→pass | 3,770 | 2,720 | -28% | 1 | 1 | 0% | 684 | 1,114 | +63% | 0 | 0 | — |
case-15 | fail→pass | 9,113 | 5,269 | -42% | 1 | 1 | 0% | 1,678 | 1,629 | -3% | 0 | 0 | — |
case-16 | pass→pass | 15,373 | 4,037 | -74% | 1 | 1 | 0% | 3,486 | 1,488 | -57% | 0 | 0 | — |
case-17 | pass→fail | 9,391 | 1,979 | -79% | 1 | 1 | 0% | 1,419 | 1,047 | -26% | 0 | 0 | — |
case-18 | pass→pass | 10,149 | 3,422 | -66% | 1 | 1 | 0% | 1,723 | 1,230 | -29% | 0 | 0 | — |
case-19 | fail→pass | 12,023 | 8,767 | -27% | 1 | 1 | 0% | 2,074 | 2,228 | +7% | 0 | 0 | — |
case-20 | pass→pass | 10,744 | 5,718 | -47% | 1 | 1 | 0% | 1,846 | 1,651 | -11% | 0 | 0 | — |
case-21 | fail→pass | 10,113 | 2,529 | -75% | 1 | 1 | 0% | 1,916 | 1,130 | -41% | 0 | 0 | — |
case-22 | pass→pass | 3,970 | 3,227 | -19% | 1 | 1 | 0% | 722 | 1,276 | +77% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +41 percentage points is the difference between those two pass rates over the 22 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.