Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Scaffold a local Codewhale plugin bundle with a versioned manifest, namespaced Skills, and an explicit trust review.
.claude/skills/hmbown-plugin-creator/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | -20% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 6% | 0% |
| case-03 | ✗→✓ | ▲ Improved | -16% | 0% |
| case-04 | ✗→✓ | ▲ Improved | -28% | 0% |
| case-05 | ✗→✓ | ▲ Improved | -21% | 0% |
Use this skill when a user wants a local Codewhale plugin bundle. Trusted and enabled bundles may add declarative Skills, commands, agents, hooks, and MCP servers (stdio and remote) through the existing engines. LSP, native extensions, filesystem roots, and lifecycle mutation are inventory-only.
~/.codewhale/plugins/<plugin-name>/<workspace>/.codewhale/plugins/<plugin-name>/plugin.toml:tomlschema_version = 1 [plugin] name = "my-plugin" version = "0.1.0" description = "What this bundle provides" [skills] path = "skills"
skills/<skill-name>/SKILL.md. Codewhale exposes itas my-plugin:<skill-name>, never as an unqualified command.
[mcp_servers.<name>] only when the bundle needs an existing MCPengine. Keep stdio commands and paths inside the bundle. Map local environment values only as exact ${SOURCE_ENV} references. For remote MCP, use HTTPS (or loopback HTTP), forbid URL user information/query/fragment, use only environment-backed headers or bearer tokens, and declare the exact normalized endpoint host set in [capabilities].network_hosts. Never place credentials in the manifest.
commands/*.md), agents (agents/*.toml), and hooks(hooks/*.toml) activate under the current policy — workspace bundles win same-name collisions over user and built-in bundles. LSP, native extensions, filesystem roots, and lifecycle mutation are inventory-only: declare them only when inventorying future work. A bundle that declares only unsupported surfaces cannot be enabled.
/plugin validate <plugin-name>/plugin show <plugin-name>/plugin enable <plugin-name> to open the content/capability review/plugin trust ... confirmation shown, then enable again/skills inspect reports plugin provenance and /plugin listreports the expected trust and activation state. Trust stages the reviewed content but does not activate it. After enablement, follow the host's reload notice: use /reload or a new session to apply changes to a live session's pinned skills and tools.
Every user and workspace bundle starts untrusted and disabled. Reuse the existing /plugin marketplace, install, update, review and reload surfaces; do not add a parallel installer, registry or automatic trust flow. Catalog membership alone never installs, trusts or enables a plugin.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 19,897 | 44,812 | +125% | 1 | 1 | 0% | 3,448 | 2,745 | -20% | 0 | 0 | — |
case-02 | fail→pass | 15,032 | 12,660 | -16% | 1 | 1 | 0% | 2,710 | 2,866 | +6% | 0 | 0 | — |
case-03 | fail→pass | 20,964 | 14,346 | -32% | 1 | 1 | 0% | 3,499 | 2,930 | -16% | 0 | 0 | — |
case-04 | fail→pass | 14,349 | 11,357 | -21% | 1 | 1 | 0% | 1,952 | 1,411 | -28% | 0 | 0 | — |
case-05 | fail→pass | 9,322 | 4,016 | -57% | 1 | 1 | 0% | 1,459 | 1,153 | -21% | 0 | 0 | — |
case-06 | fail→pass | 9,536 | 10,325 | +8% | 1 | 1 | 0% | 1,472 | 1,397 | -5% | 0 | 0 | — |
case-07 | fail→pass | 10,388 | 9,363 | -10% | 1 | 1 | 0% | 1,639 | 1,844 | +13% | 0 | 0 | — |
case-08 | fail→pass | 16,788 | 8,085 | -52% | 1 | 1 | 0% | 2,648 | 1,905 | -28% | 0 | 0 | — |
case-09 | fail→pass | 15,850 | 9,134 | -42% | 1 | 1 | 0% | 2,185 | 1,996 | -9% | 0 | 0 | — |
case-10 | pass→pass | 10,168 | 2,810 | -72% | 1 | 1 | 0% | 777 | 966 | +24% | 0 | 0 | — |
case-11 | fail→pass | 13,091 | 5,582 | -57% | 1 | 1 | 0% | 1,688 | 1,553 | -8% | 0 | 0 | — |
case-12 | fail→pass | 11,875 | 3,086 | -74% | 1 | 1 | 0% | 1,714 | 1,006 | -41% | 0 | 0 | — |
case-13 | fail→pass | 10,366 | 4,190 | -60% | 1 | 1 | 0% | 1,269 | 1,263 | -0% | 0 | 0 | — |
case-14 | pass→pass | 9,793 | 4,278 | -56% | 1 | 1 | 0% | 1,441 | 1,137 | -21% | 0 | 0 | — |
case-15 | fail→pass | 17,305 | 7,059 | -59% | 1 | 1 | 0% | 2,292 | 1,621 | -29% | 0 | 0 | — |
case-16 | pass→pass | 11,031 | 3,704 | -66% | 1 | 1 | 0% | 1,337 | 1,062 | -21% | 0 | 0 | — |
case-17 | fail→pass | 28,957 | 5,462 | -81% | 1 | 1 | 0% | 2,903 | 1,025 | -65% | 0 | 0 | — |
case-18 | pass→pass | 12,439 | 3,484 | -72% | 1 | 1 | 0% | 1,469 | 1,020 | -31% | 0 | 0 | — |
case-19 | fail→fail | 9,640 | 3,941 | -59% | 1 | 1 | 0% | 1,489 | 1,126 | -24% | 0 | 0 | — |
case-20 | pass→pass | 11,663 | 7,477 | -36% | 1 | 1 | 0% | 1,754 | 1,277 | -27% | 0 | 0 | — |
case-21 | fail→pass | 22,555 | 14,383 | -36% | 1 | 1 | 0% | 3,763 | 2,291 | -39% | 0 | 0 | — |
case-22 | fail→pass | 17,184 | 9,290 | -46% | 1 | 1 | 0% | 1,421 | 2,021 | +42% | 0 | 0 | — |
case-23 | fail→fail | 9,827 | 10,903 | +11% | 1 | 1 | 0% | 1,461 | 2,395 | +64% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of +70 percentage points is the difference between those two pass rates over the 23 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
| Model | Method | Date | Lift |
|---|---|---|---|
| gemini-3.6-flash | verified | 9/19/2026 | +57% |
| gemini-3.6-flash | verified | 8/17/2026 | +73% |
| gemini-3.6-flash | verified | 8/7/2026 | +59% |
Other measured skills in the registry, with their headline benchmark lift.