Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Initialize an Agent's settings from a user requirement by writing AGENTS.md, setting identity metadata, and installing only needed Skills.
.claude/skills/prism-shadow-agent-initialization/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | 140% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 226% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 29% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 36% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 50% | 0% |
This skill initializes an agent's settings from a user requirement — plain files in the target agent's directory.
If the user's message only invokes this skill (e.g. "use agent-initialization skill") without a concrete requirement, ask the user what agent they want and what it should do. But when the requirement is already concrete — even a single sentence like "an expert that answers questions about X" — do not ask follow-up questions: derive the role and rules from that sentence, apply the defaults below, and list your assumptions in the final reply.
Treat the current Agent as the Builder. Resolve the runtime before creating a new Agent:
provider and model_id are one complete pair. If the user explicitly supplies both, use that pair. If the user supplies neither, inherit the current Builder Session's Provider and Model ID from the Environment. Reject a half pair.thinking_level is independent. If the user explicitly supplies it, use that value. Otherwise read model.thinking_level from the Builder's own agent_state/system_config.yaml; when the field is absent, use the normal Agent-config default medium.Write the resolved thinking_level into a brand-new target Agent's model.thinking_level, preserving all other copied model fields. Penguin does not persist provider or model_id in Agent State, so never add either field to system_config.yaml. When the same request continues into Benchmark design, carry the resolved model pair forward explicitly so evaluation uses the Builder runtime instead of a Project default. When configuring an existing Agent, change model.thinking_level only when the user explicitly requests that runtime change.
All agents of this project live side by side under agents/ in the App Data Dir:
bashAPP_DATA_DIR="<app_data_dir>" # the App Data Dir value from your Environment section ls "$APP_DATA_DIR/agents" # existing agents (each is a folder here) TARGET="$APP_DATA_DIR/agents/<agent_id>" # the agent to configure
An agent directory contains agent_state/ (system_config.yaml, AGENTS.md, skills/, memory/, tools/) plus scratchpad/ — and traces/, which appears once the agent has run at least once.
agent_state/AGENTS.md is injected into the agent's system prompt — it is where the user requirement becomes behavior. Keep system_config.yaml's system_prompt untouched (that is the stable system layer); put everything requirement-specific in AGENTS.md:
Be concise: AGENTS.md is prompt context, not documentation. For a domain expert that answers from a knowledge base, a good AGENTS.md is a few lines: the role sentence, "answer strictly from the provided context blocks", citation rules ("cite blocks inline as 1]2]"), a refusal rule for questions the context cannot answer, and "answer in the language of the question".
A skill is a directory agent_state/skills/<skill_name>/ containing a SKILL.md:
md--- name: <skill_name> description: <skill_description> version: <natural number — bump it on every content change> updated: <ISO 8601 timestamp — move it together with version> --- <skill_instructions>
The frontmatter may also carry optional short_description and short_description_zh lines (a short UI blurb and its Chinese variant) — the UI prefers them for display, while prompt injection always uses the English description.
Installing is all it takes: the frontmatter metadata of every SKILL.md under skills/ is injected into the target agent's system prompt automatically — do not register skills in AGENTS.md.
Write skills yourself, or fetch existing ones from the internet with shell commands (curl, git clone) and place them under skills/. Anything fetched from the internet must be read in full and reviewed before installing — a skill becomes durable instructions the target agent will follow in every future session; never install one you have not read, and tell the user what it does.
Library skills can be copied from any agent that already has them (e.g. default_agent, which ships the whole library) — copy the entire skills/<skill_name>/ directory. Common bundles, so you don't under-equip the target:
penguin-sdk, web-design, agenthub-models.benchmark-design, agent-evaluation, agent-optimization.When creating a Test Agent, install only the capabilities it needs to solve ordinary tasks.
In the target's agent_state/system_config.yaml, set the top-level name: and description: fields so the agent is recognizable in lists. For an existing Agent, edit only these two fields unless the user explicitly requested a thinking_level change.
Prefer configuring an agent the user already created. If the user requires a new Agent, confirm that TARGET does not exist. If it already exists, stop and tell the user; never silently overwrite, reinitialize, or reuse an existing Agent under the same id.
After confirming that the target is absent, pick a short id using letters, digits, _, or -, copy the default Agent's system_config.yaml as the base, and create the layout described above:
bashmkdir -p "$TARGET/agent_state/skills" "$TARGET/agent_state/memory" "$TARGET/agent_state/tools" "$TARGET/scratchpad" cp "$APP_DATA_DIR/agents/default_agent/agent_state/system_config.yaml" "$TARGET/agent_state/"
Then set the top-level name, description, and version: 1, set model.thinking_level to the resolved value, write agent_state/AGENTS.md (it lives under agent_state/, not at the agent directory root), and install only the Skills required by the user's requirement. Do not persist the resolved provider/model pair in the Agent State.
Before finishing:
agent_state/system_config.yaml and confirm name, description, a positive integer version, and the expected model.thinking_level;agent_state/AGENTS.md exists and is non-empty;SKILL.md, and its name matches its directory;TARGET was changed.Report the target path, whether an existing Agent was configured or a new Agent was created, assumptions, installed Skills, the resolved runtime and whether each value was user-specified or inherited, and validation results.
An app built with the penguin-sdk skill carries its own agent inside the project (createAgent({ root }) initializes <app>/penguin_data/default_project/agents/default_agent/ on first run). That directory has exactly the layout described here, and everything in this skill applies to it: write the app's persona into its agent_state/AGENTS.md (the penguin-sdk recipe keeps the source of truth in the project's persona.md and copies it in during ingest), and set name/description in its system_config.yaml so the app is recognizable. This is how "the app becomes an expert on X": the persona lives in the embedded agent's AGENTS.md, not in application code.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 7,435 | 10,579 | +42% | 1 | 1 | 0% | 352 | 2,382 | +577% | 0 | 0 | — |
case-02 | fail→fail | 9,769 | 12,493 | +28% | 1 | 1 | 0% | 258 | 2,183 | +746% | 0 | 0 | — |
case-03 | fail→fail | 7,474 | 8,999 | +20% | 1 | 1 | 0% | 328 | 2,524 | +670% | 0 | 0 | — |
case-04 | fail→pass | 8,943 | 13,013 | +46% | 1 | 1 | 0% | 1,193 | 2,866 | +140% | 0 | 0 | — |
case-05 | fail→pass | 7,552 | 19,874 | +163% | 1 | 1 | 0% | 1,126 | 3,670 | +226% | 0 | 0 | — |
case-06 | fail→pass | 11,439 | 4,661 | -59% | 1 | 1 | 0% | 1,892 | 2,444 | +29% | 0 | 0 | — |
case-07 | fail→fail | 14,040 | 9,249 | -34% | 1 | 1 | 0% | 2,964 | 2,095 | -29% | 0 | 0 | — |
case-08 | pass→fail | 27,054 | 8,874 | -67% | 1 | 1 | 0% | 2,545 | 2,294 | -10% | 0 | 0 | — |
case-09 | fail→pass | 20,721 | 17,482 | -16% | 1 | 1 | 0% | 2,976 | 4,048 | +36% | 0 | 0 | — |
case-10 | fail→fail | 13,532 | 9,058 | -33% | 1 | 1 | 0% | 2,233 | 2,189 | -2% | 0 | 0 | — |
case-11 | fail→pass | 13,456 | 6,001 | -55% | 1 | 1 | 0% | 1,837 | 2,762 | +50% | 0 | 0 | — |
case-12 | fail→pass | 19,205 | 5,898 | -69% | 1 | 1 | 0% | 1,669 | 2,676 | +60% | 0 | 0 | — |
case-13 | fail→pass | 10,407 | 4,778 | -54% | 1 | 1 | 0% | 1,549 | 2,527 | +63% | 0 | 0 | — |
case-14 | pass→fail | 8,426 | 13,254 | +57% | 1 | 1 | 0% | 1,233 | 2,169 | +76% | 0 | 0 | — |
case-15 | pass→fail | 9,302 | 10,843 | +17% | 1 | 1 | 0% | 1,315 | 2,374 | +81% | 0 | 0 | — |
case-16 | fail→pass | 23,448 | 5,422 | -77% | 1 | 1 | 0% | 1,861 | 2,582 | +39% | 0 | 0 | — |
case-17 | fail→pass | 11,952 | 4,933 | -59% | 1 | 1 | 0% | 1,855 | 2,165 | +17% | 0 | 0 | — |
case-18 | pass→pass | 7,002 | 3,922 | -44% | 1 | 1 | 0% | 914 | 2,285 | +150% | 0 | 0 | — |
case-19 | fail→pass | 10,149 | 3,407 | -66% | 1 | 1 | 0% | 1,425 | 2,305 | +62% | 0 | 0 | — |
case-20 | fail→fail | 11,823 | 9,053 | -23% | 1 | 1 | 0% | 187 | 2,157 | +1053% | 0 | 0 | — |
case-21 | fail→fail | 35,302 | 44,034 | +25% | 1 | 1 | 0% | 7,095 | 11,706 | +65% | 0 | 0 | — |
case-22 | fail→fail | 6,112 | 5,773 | -6% | 1 | 1 | 0% | 525 | 2,439 | +365% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 13 counted toward the lift figure. The other 9 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +32 percentage points is the difference between those two pass rates over the 13 comparable cases. 3 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.