Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Wrap any command-line tool into a JSON-emitting agent skill. Use when you want to make a CLI reliably callable and parseable by an AI agent — introspect its --help, define a structured-output (--json) contract and a stable error contract, then generate a SKILL.md wrapper with verified examples. Triggers on "wrap this CLI", "make X agent-usable", "generate a skill for this command", "turn a tool into an agent skill".
.claude/skills/coco-research-cli-anything/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | -23% | 0% |
| case-03 | ✗→✓ | ▲ Improved | -24% | 0% |
| case-09 | ✗→✓ | ▲ Improved | -1% | 0% |
| case-11 | ✗→✓ | ▲ Improved | -3% | 0% |
| case-15 | ✗→✓ | ▲ Improved | 15% | 0% |
<!-- Methodology adapted from HKUDS/CLI-Anything (https://github.com/HKUDS/CLI-Anything) — Apache-2.0. -->
Turn any command-line tool into something an AI agent can call reliably and parse. The output is a thin, well-documented wrapper contract — never a reimplementation of the tool. Use the real tool; document how to drive it.
You have a CLI (your own, or a third-party binary) and you want an agent to invoke it predictably, read structured results, and recover from errors. This skill produces a SKILL.md wrapper that encodes that contract.
Enumerate the real surface before writing anything:
<tool> --help and <tool> <subcommand> --help for each subcommand.--help is thin, read the source or man page. Every claim in the wrapper must trace to observed behavior.--json machine-readable mode for every command an agent will call. Document the exact JSON keys the agent should read (one shape per command).{"error": "...", "code": N} if the tool supports it, otherwise the captured stderr) — never a bare stack trace.Standard structure so an agent can discover and drive the tool:
name, description (include natural trigger phrases).Commands — each command: purpose, exact invocation, args, and the result shape it returns.Examples — 2-4 real invocations with expected output.Errors — the exit-code/error contract from Step 3.Notes — auth, side effects, idempotency, network egress, prerequisites.--json shape parses.skills/coco-cli/SKILL.md is a wrapper produced with this pattern over the cocosuperintelligence command — note how it documents the (human-only) output contract honestly rather than inventing a JSON mode the tool lacks.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 25,129 | 28,438 | +13% | 1 | 1 | 0% | 4,205 | 4,629 | +10% | 0 | 0 | — |
case-02 | fail→pass | 30,297 | 22,021 | -27% | 1 | 1 | 0% | 5,288 | 4,067 | -23% | 0 | 0 | — |
case-03 | fail→pass | 31,390 | 27,008 | -14% | 1 | 1 | 0% | 6,283 | 4,766 | -24% | 0 | 0 | — |
case-04 | pass→pass | 12,829 | 12,945 | +1% | 1 | 1 | 0% | 1,936 | 1,856 | -4% | 0 | 0 | — |
case-05 | pass→pass | 12,402 | 5,883 | -53% | 1 | 1 | 0% | 2,118 | 1,530 | -28% | 0 | 0 | — |
case-06 | pass→pass | 16,531 | 14,401 | -13% | 1 | 1 | 0% | 2,585 | 2,561 | -1% | 0 | 0 | — |
case-07 | pass→pass | 13,796 | 12,571 | -9% | 1 | 1 | 0% | 1,682 | 1,935 | +15% | 0 | 0 | — |
case-08 | fail→fail | 19,721 | 7,977 | -60% | 1 | 1 | 0% | 2,117 | 1,841 | -13% | 0 | 0 | — |
case-09 | fail→pass | 17,696 | 16,277 | -8% | 1 | 1 | 0% | 2,422 | 2,401 | -1% | 0 | 0 | — |
case-10 | pass→pass | 16,918 | 10,718 | -37% | 1 | 1 | 0% | 1,790 | 1,668 | -7% | 0 | 0 | — |
case-11 | fail→pass | 17,802 | 8,551 | -52% | 1 | 1 | 0% | 2,061 | 1,993 | -3% | 0 | 0 | — |
case-12 | pass→pass | 17,271 | 5,003 | -71% | 1 | 1 | 0% | 1,878 | 1,428 | -24% | 0 | 0 | — |
case-13 | pass→pass | 20,038 | 14,324 | -29% | 1 | 1 | 0% | 2,351 | 2,296 | -2% | 0 | 0 | — |
case-14 | pass→pass | 18,396 | 8,722 | -53% | 1 | 1 | 0% | 2,319 | 2,034 | -12% | 0 | 0 | — |
case-15 | fail→pass | 16,819 | 17,914 | +7% | 1 | 1 | 0% | 2,653 | 3,056 | +15% | 0 | 0 | — |
case-16 | pass→pass | 15,420 | 12,240 | -21% | 1 | 1 | 0% | 1,528 | 1,704 | +12% | 0 | 0 | — |
case-17 | pass→pass | 13,163 | 12,268 | -7% | 1 | 1 | 0% | 2,144 | 1,766 | -18% | 0 | 0 | — |
case-18 | fail→pass | 16,040 | 9,314 | -42% | 1 | 1 | 0% | 2,025 | 2,132 | +5% | 0 | 0 | — |
case-19 | pass→pass | 23,532 | 16,984 | -28% | 1 | 1 | 0% | 2,944 | 2,465 | -16% | 0 | 0 | — |
case-20 | pass→pass | 27,216 | 39,457 | +45% | 1 | 1 | 0% | 4,675 | 8,945 | +91% | 0 | 0 | — |
case-21 | pass→pass | 21,081 | 27,967 | +33% | 1 | 1 | 0% | 2,703 | 4,562 | +69% | 0 | 0 | — |
case-22 | pass→pass | 17,992 | 22,444 | +25% | 1 | 1 | 0% | 3,553 | 4,403 | +24% | 0 | 0 | — |
case-23 | pass→pass | 20,458 | 16,147 | -21% | 1 | 1 | 0% | 2,526 | 2,474 | -2% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of +26 percentage points is the difference between those two pass rates over the 23 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.