Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Manage model API keys, default models and per-agent vault secrets with the penguin CLI.
.claude/skills/prism-shadow-penguin-cli/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | -3% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 83% | 0% |
| case-03 | ✗→✓ | ▲ Improved | -58% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 38% | 0% |
| case-05 | ✗→✓ | ▲ Improved | -37% | 0% |
The penguin CLI manages model credentials, default models and per-agent vault secrets. Its primary job is model configuration: penguin config model add registers a model and penguin config model list shows the models currently available. Configuration goes through the CLI only — never read or hand-edit the underlying hidden files.
If the user's message only invokes this skill (e.g. "use penguin-cli skill") without a concrete request, ask the user what they want to configure. Do not run any command until the goal is clear.
Add or update a model (upsert by the (provider, model_id) pair; re-run with more options to amend an entry):
bashpenguin config model add --provider <group> --model-id <upstream_id> [--api-key <key>] [--base-url <url>] \ [--client-type <type>] [--context-window <n>] [--max-tokens <n>] [--vision | --no-vision] \ [--price-cache-read <n>] [--price-cache-write <n>] [--price-output <n>] \ [--project-id <id>] [--root <dir>] [--set-default]
(provider, model_id) pair, so --provider and --model-id are both required — the group is never inferred from the model id, because gateways resell vendor models under their upstream ids and a wrong guess would send the key to another vendor's endpoint. --model-id takes the provider's upstream model id (what the API expects) and is persisted as the entry's request id, so it reaches the API unchanged; --provider names the group (deepseek, openai, anthropic, google, openrouter, siliconflow, … — custom for any other endpoint).--client-type openai --base-url <endpoint>; omit --client-type to auto-route by model id.--vision / --no-vision mark whether the model accepts images; omitting both keeps the current value (default is vision-capable).--max-tokens <n> pins a per-model output cap (positive integer), overriding the Agent's model.max_tokens; omit to inherit. Lower it for small-context models — the per-Agent default (32000) cannot fit into e.g. a 32k context window together with any prompt.penguin config model ... and penguin config vault ... commands accept --root <dir> to target another data root (default PENGUIN_HOME, then ~/.penguin/data). Two configuration targets — treat the difference as a hard rule:--root is correct.--root must point at the app's own data directory inside the project (e.g. --root ./penguin_data, the same path the app gives createAgent({ root })) unless the user explicitly chose another location — never write an app's models or keys into the global ~/.penguin/data, which belongs to the person running Penguin, not to the app.penguin config model list --root <app root> should show the app's entries, and the global list (no --root) should stay clean.Other model commands:
bashpenguin config model default --model-id <upstream_id> --provider <group> [--root <dir>] # set the project default model penguin config model vision --model-id <upstream_id> --provider <group> [--root <dir>] # set the project vision model (reads images for text-only sessions) penguin config model list [--root <dir>] # list models; api_key is shown masked
The vault holds an agent's environment-variable secrets (third-party API keys etc.); values are injected into that agent's shell subprocesses:
bashpenguin config vault set --key <NAME> --value <value> [--project-id <id>] [--agent-id <id>] [--root <dir>] penguin config vault list [--project-id <id>] [--agent-id <id>] [--root <dir>] # values are shown masked penguin config vault remove --key <NAME> [--project-id <id>] [--agent-id <id>] [--root <dir>]
--project-id defaults to default_project, --agent-id to default_agent.bashpenguin config lang <en|zh> # persist the CLI language via PENGUIN_LANG in your shell rc
penguin run -m "<task>" [--provider <group> --model-id <id>] [--agent-id <id>] [--workspace <path>] [--approve <mode>] runs one task; penguin chat [--resume [session_id]] starts or resumes an interactive chat with the same options. The model reference stays a pair here too: pass --provider and --model-id together, or neither to run on the project's default model — one without the other is rejected.
Paths use <app_data_dir>, the App Data Dir value from your Environment section.
<app_data_dir>/.project_config.toml — the project's single hidden config file: model list, settings and per-model credentials (api_key etc. inlined in each model entry). Configuration is CLI-only — never read, print or hand-edit this file.<app_data_dir>/agents/<agent_id>/agent_state/.vault.toml — that agent's vault entries, hidden file; same rule, manage it with penguin config vault.| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 12,898 | 8,721 | -32% | 1 | 1 | 0% | 2,231 | 2,174 | -3% | 0 | 0 | — |
case-02 | fail→pass | 10,651 | 7,197 | -32% | 1 | 1 | 0% | 992 | 1,812 | +83% | 0 | 0 | — |
case-03 | fail→pass | 27,804 | 2,404 | -91% | 1 | 1 | 0% | 4,253 | 1,782 | -58% | 0 | 0 | — |
case-04 | fail→pass | 13,742 | 3,766 | -73% | 1 | 1 | 0% | 1,483 | 2,040 | +38% | 0 | 0 | — |
case-05 | fail→pass | 16,520 | 8,884 | -46% | 1 | 1 | 0% | 3,108 | 1,970 | -37% | 0 | 0 | — |
case-06 | fail→pass | 7,230 | 2,293 | -68% | 1 | 1 | 0% | 1,241 | 1,813 | +46% | 0 | 0 | — |
case-07 | fail→pass | 38,795 | 7,751 | -80% | 1 | 1 | 0% | 6,327 | 1,939 | -69% | 0 | 0 | — |
case-08 | fail→pass | 26,256 | 2,160 | -92% | 1 | 1 | 0% | 4,677 | 1,766 | -62% | 0 | 0 | — |
case-09 | fail→pass | 29,044 | 2,016 | -93% | 1 | 1 | 0% | 4,308 | 1,724 | -60% | 0 | 0 | — |
case-10 | fail→pass | 34,209 | 8,274 | -76% | 1 | 1 | 0% | 6,940 | 2,127 | -69% | 0 | 0 | — |
case-11 | fail→pass | 22,253 | 2,029 | -91% | 1 | 1 | 0% | 4,147 | 1,717 | -59% | 0 | 0 | — |
case-12 | fail→pass | 15,111 | 6,694 | -56% | 1 | 1 | 0% | 2,629 | 1,725 | -34% | 0 | 0 | — |
case-13 | fail→pass | 13,948 | 8,219 | -41% | 1 | 1 | 0% | 1,612 | 1,875 | +16% | 0 | 0 | — |
case-14 | fail→pass | 8,225 | 2,753 | -67% | 1 | 1 | 0% | 1,457 | 1,772 | +22% | 0 | 0 | — |
case-15 | pass→pass | 18,500 | 6,642 | -64% | 1 | 1 | 0% | 738 | 1,613 | +119% | 0 | 0 | — |
case-16 | fail→pass | 15,315 | 10,357 | -32% | 1 | 1 | 0% | 1,897 | 2,417 | +27% | 0 | 0 | — |
case-17 | fail→pass | 20,943 | 7,532 | -64% | 1 | 1 | 0% | 2,873 | 1,765 | -39% | 0 | 0 | — |
case-18 | fail→pass | 29,806 | 8,450 | -72% | 1 | 1 | 0% | 5,033 | 2,093 | -58% | 0 | 0 | — |
case-19 | fail→pass | 22,737 | 7,789 | -66% | 1 | 1 | 0% | 3,533 | 1,970 | -44% | 0 | 0 | — |
case-20 | fail→pass | 26,842 | 7,932 | -70% | 1 | 1 | 0% | 3,978 | 1,778 | -55% | 0 | 0 | — |
case-21 | fail→pass | 35,714 | 2,537 | -93% | 1 | 1 | 0% | 2,784 | 1,715 | -38% | 0 | 0 | — |
case-22 | pass→fail | 11,198 | 11,647 | +4% | 1 | 1 | 0% | 2,015 | 3,440 | +71% | 0 | 0 | — |
case-23 | pass→fail | 4,704 | 3,917 | -17% | 1 | 1 | 0% | 864 | 2,135 | +147% | 0 | 0 | — |
case-24 | pass→pass | 15,427 | 12,798 | -17% | 1 | 1 | 0% | 1,857 | 2,475 | +33% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 24 cases were attempted. The headline lift of +75 percentage points is the difference between those two pass rates over the 24 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.