Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Designs multi-agent system architectures with orchestration patterns, tool schemas, and performance evaluation. Use when building AI agent systems, designing agent workflows, creating tool schemas, or evaluating agent performance.
.claude/skills/borghei-agent-designer/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | -12% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 48% | 0% |
| case-05 | ✗→✓ | ▲ Improved | -4% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -10% | 0% |
| case-08 | ✗→✓ | ▲ Improved | -35% | 0% |
A toolkit for designing, architecting, and evaluating multi-agent systems. It provides structured approaches to agent architecture patterns, tool design principles, communication strategies, and performance evaluation frameworks for building robust, scalable AI agent systems.
Before designing the system, confirm these inputs. If any is unknown or vague, ASK — do not assume:
tool_schema_generator.py emits)Stop rule: ask only the 2-3 that most change the output. If the user says "just draft it," proceed and list your assumptions at the top of the artifact.
| Tool | Purpose | Command | |------|---------|---------| | agent_planner.py | Design architecture from requirements (pattern, roles, topology, Mermaid diagram, roadmap) | python agent_planner.py requirements.json -o my_system --format both | | agent_evaluator.py | Evaluate performance from execution logs (success, cost, latency, bottlenecks) | python agent_evaluator.py execution_logs.json -o perf_report --format both --detailed | | tool_schema_generator.py | Generate OpenAI/Anthropic tool schemas with validation | python tool_schema_generator.py tools.json -o my_tools --format both --validate |
Load the reference that matches the task — keep this file lean and pull detail on demand:
Covers:
Does NOT cover:
engineering/agent-workflow-designer for workflow execution)engineering/prompt-engineer-toolkit)engineering/mcp-server-builder)engineering/self-improving-agent)| Skill | Integration | Data Flow | |-------|-------------|-----------| | engineering/agent-workflow-designer | Workflow definitions consume architecture designs from Agent Designer | Agent roles and communication topology feed into workflow step definitions | | engineering/prompt-engineer-toolkit | System prompts are crafted per agent role defined by Agent Designer | Agent role specifications and responsibilities inform prompt structure and constraints | | engineering/mcp-server-builder | Tool schemas generated here map to MCP server tool implementations | tool_schema_generator.py output provides the schema contract that MCP servers implement | | engineering/self-improving-agent | Evaluation reports feed into self-improvement loops | agent_evaluator.py bottleneck analysis drives autonomous optimization decisions | | engineering/observability-designer | Monitoring architecture aligns with agent topology and communication links | Agent definitions and communication patterns define what to instrument and alert on | | engineering/agent-protocol | Protocol standards govern inter-agent message formats designed here | Communication topology patterns must comply with agent protocol specifications |
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-04 | fail→pass | 19,486 | 9,722 | -50% | 1 | 1 | 0% | 3,282 | 2,898 | -12% | 0 | 0 | — |
case-01 | fail→fail | 18,185 | 23,999 | +32% | 1 | 1 | 0% | 3,048 | 5,388 | +77% | 0 | 0 | — |
case-02 | fail→fail | 16,950 | 23,270 | +37% | 1 | 1 | 0% | 3,273 | 5,797 | +77% | 0 | 0 | — |
case-03 | fail→pass | 13,560 | 11,207 | -17% | 1 | 1 | 0% | 2,198 | 3,246 | +48% | 0 | 0 | — |
case-05 | fail→pass | 19,869 | 14,688 | -26% | 1 | 1 | 0% | 3,962 | 3,810 | -4% | 0 | 0 | — |
case-06 | fail→fail | 25,573 | 28,986 | +13% | 1 | 1 | 0% | 5,066 | 6,949 | +37% | 0 | 0 | — |
case-07 | fail→pass | 12,052 | 3,508 | -71% | 1 | 1 | 0% | 2,126 | 1,909 | -10% | 0 | 0 | — |
case-08 | fail→pass | 16,821 | 3,734 | -78% | 1 | 1 | 0% | 3,012 | 1,954 | -35% | 0 | 0 | — |
case-09 | fail→fail | 14,967 | 19,193 | +28% | 1 | 1 | 0% | 2,496 | 4,611 | +85% | 0 | 0 | — |
case-10 | pass→pass | 14,583 | 11,763 | -19% | 1 | 1 | 0% | 2,271 | 3,290 | +45% | 0 | 0 | — |
case-11 | pass→pass | 12,802 | 12,288 | -4% | 1 | 1 | 0% | 1,999 | 3,346 | +67% | 0 | 0 | — |
case-12 | pass→pass | 7,097 | 9,338 | +32% | 1 | 1 | 0% | 1,096 | 2,725 | +149% | 0 | 0 | — |
case-13 | pass→pass | 6,049 | 8,066 | +33% | 1 | 1 | 0% | 978 | 2,680 | +174% | 0 | 0 | — |
case-14 | pass→pass | 7,104 | 8,370 | +18% | 1 | 1 | 0% | 1,132 | 2,626 | +132% | 0 | 0 | — |
case-15 | pass→pass | 8,557 | 8,596 | +0% | 1 | 1 | 0% | 1,381 | 2,730 | +98% | 0 | 0 | — |
case-16 | pass→pass | 15,563 | 15,017 | -4% | 1 | 1 | 0% | 2,459 | 3,612 | +47% | 0 | 0 | — |
case-17 | pass→pass | 10,012 | 12,426 | +24% | 1 | 1 | 0% | 1,692 | 3,474 | +105% | 0 | 0 | — |
case-18 | pass→pass | 6,193 | 7,043 | +14% | 1 | 1 | 0% | 1,001 | 2,491 | +149% | 0 | 0 | — |
case-19 | pass→pass | 10,208 | 15,055 | +47% | 1 | 1 | 0% | 1,761 | 3,968 | +125% | 0 | 0 | — |
case-20 | fail→pass | 20,953 | 3,293 | -84% | 1 | 1 | 0% | 3,447 | 1,858 | -46% | 0 | 0 | — |
case-21 | fail→pass | 9,165 | 1,658 | -82% | 1 | 1 | 0% | 1,432 | 1,580 | +10% | 0 | 0 | — |
case-22 | fail→pass | 10,102 | 2,310 | -77% | 1 | 1 | 0% | 1,557 | 1,691 | +9% | 0 | 0 | — |
case-23 | pass→pass | 16,757 | 14,486 | -14% | 1 | 1 | 0% | 2,684 | 3,538 | +32% | 0 | 0 | — |
case-24 | pass→pass | 12,983 | 11,487 | -12% | 1 | 1 | 0% | 2,308 | 3,436 | +49% | 0 | 0 | — |
case-25 | fail→pass | 13,779 | 8,134 | -41% | 1 | 1 | 0% | 2,168 | 2,621 | +21% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 25 cases were attempted. The headline lift of +36 percentage points is the difference between those two pass rates over the 25 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.