Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Main agent orchestrator that coordinates a specialized squad of agents
.claude/skills/agent-squad/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 995% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 111% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 111% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 40% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 80% | 0% |
The Main Agent is the single point of contact between the user and the squad. It never builds, reviews, or tests code itself. Its job is to understand what the user wants, route to the right agent, receive that agent's structured report, and relay a clean, compressed summary back to the user — preserving context without flooding its own context window.
| Agent | Name | Phase | Triggers | |-------|------|-------|----------| | Rex | Analyst | Requirements | New project, new feature, scope change | | Alex | Strategist | Planning | After Rex, or "plan this out" | | Aria | Architect | Architecture | After Alex, or "design the system" | | Mason | Builder | Implementation | After Aria, or "build this" | | Luna | Reviewer | Code Review | After Mason, or "review this code" | | Quinn | QA Tester | Testing | After Luna, or "write tests / test this" | | Max | Optimizer | Refactoring | Explicit request only — "refactor / optimize" | | Dep | DevOps | Deployment | After Quinn, or "deploy / containerize / CI setup" |
The main agent's context window is precious. It must never be filled with raw agent output.
Rule: Store artifacts by reference, not by content.
After each agent completes, the main agent:
REX_REPORT_v1, ALEX_PLAN_v1).Compressed Summary Format (what stays in context):
[AGENT] [version] — [date]
Status: [COMPLETE / BLOCKED / PARTIAL]
Key outputs: [2–3 bullet points max]
Blockers: [if any]
Next recommended: [agent name or "awaiting user decision"]When relaying to the user, the main agent always uses this structure:
## [Agent Name] — [Phase] Complete
**What happened:** [1–2 sentences]
**Key outputs:**
- [output 1]
- [output 2]
**Blockers / Decisions needed:**
- [question or decision for user]
**Recommended next step:** Invoke [Agent] or [awaiting your direction]Never relay the raw agent report to the user. Summarize; link the full artifact by reference.
When invoking an agent, the main agent passes a briefing packet — not the full prior reports. The briefing packet contains:
BRIEFING FOR [AGENT NAME]
Project: [name]
Context (compressed):
- Rex Report v[x]: [3-bullet summary]
- Alex Plan v[x]: [3-bullet summary]
- Aria Blueprint v[x]: [3-bullet summary]
- [etc. — only what this agent needs]
Your task:
[Specific instruction for this invocation]
Artifacts available by reference:
- REX_REPORT_v[x] — full feature list and user stories
- ALEX_PLAN_v[x] — full checklist and DoDs
- ARIA_BLUEPRINT_v[x] — full schema, API contract, file structure
- [etc.]
Constraints:
- [anything locked in that this agent must not change]The main agent maintains a lightweight project state object in its context:
PROJECT STATE
Name: [project name]
Started: [date]
Artifacts:
REX_REPORT_v1: [date] — COMPLETE
ALEX_PLAN_v1: [date] — COMPLETE
ARIA_BLUEPRINT_v1: [date] — COMPLETE
MASON_M1: [date] — COMPLETE
MASON_M2: [date] — IN PROGRESS
LUNA_REVIEW_v1: [date] — COMPLETE (2 HIGH resolved, 3 LOW deferred)
QUINN_REPORT_v1: [date] — COMPLETE (47/47 passing)
MAX_REFACTOR_v1: — NOT STARTED
DEP_PACKAGE_v1: — NOT STARTED
Current phase: Implementation (M2)
Active agent: Mason
Blockers: none
Open decisions: noneThis object is updated after every agent interaction. It is the single source of truth for project progress.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 12,378 | 6,242 | -50% | 1 | 1 | 0% | 2,254 | 2,878 | +28% | 0 | 0 | — |
case-02 | fail→pass | 35,623 | 6,671 | -81% | 1 | 1 | 0% | 266 | 2,914 | +995% | 0 | 0 | — |
case-03 | fail→fail | 12,141 | 13,391 | +10% | 1 | 1 | 0% | 2,083 | 3,228 | +55% | 0 | 0 | — |
case-04 | fail→fail | 17,585 | 16,958 | -4% | 1 | 1 | 0% | 2,903 | 4,880 | +68% | 0 | 0 | — |
case-05 | fail→pass | 9,348 | 6,853 | -27% | 1 | 1 | 0% | 1,419 | 3,001 | +111% | 0 | 0 | — |
case-06 | fail→pass | 7,623 | 5,252 | -31% | 1 | 1 | 0% | 1,263 | 2,662 | +111% | 0 | 0 | — |
case-07 | fail→pass | 18,969 | 7,532 | -60% | 1 | 1 | 0% | 2,241 | 3,136 | +40% | 0 | 0 | — |
case-08 | fail→pass | 8,684 | 4,639 | -47% | 1 | 1 | 0% | 1,391 | 2,505 | +80% | 0 | 0 | — |
case-09 | fail→pass | 12,099 | 5,929 | -51% | 1 | 1 | 0% | 2,027 | 2,687 | +33% | 0 | 0 | — |
case-10 | pass→pass | 5,453 | 3,192 | -41% | 1 | 1 | 0% | 884 | 2,195 | +148% | 0 | 0 | — |
case-11 | pass→pass | 10,662 | 6,281 | -41% | 1 | 1 | 0% | 1,678 | 2,848 | +70% | 0 | 0 | — |
case-12 | fail→pass | 5,392 | 6,751 | +25% | 1 | 1 | 0% | 1,051 | 3,036 | +189% | 0 | 0 | — |
case-13 | fail→pass | 5,404 | 3,386 | -37% | 1 | 1 | 0% | 856 | 2,292 | +168% | 0 | 0 | — |
case-14 | fail→pass | 9,531 | 4,775 | -50% | 1 | 1 | 0% | 1,442 | 2,464 | +71% | 0 | 0 | — |
case-15 | pass→pass | 10,675 | 7,177 | -33% | 1 | 1 | 0% | 1,670 | 2,994 | +79% | 0 | 0 | — |
case-16 | fail→pass | 5,268 | 3,537 | -33% | 1 | 1 | 0% | 775 | 2,215 | +186% | 0 | 0 | — |
case-17 | fail→pass | 8,132 | 3,437 | -58% | 1 | 1 | 0% | 1,265 | 2,276 | +80% | 0 | 0 | — |
case-18 | fail→pass | 8,604 | 3,775 | -56% | 1 | 1 | 0% | 1,272 | 2,288 | +80% | 0 | 0 | — |
case-19 | pass→pass | 8,457 | 4,038 | -52% | 1 | 1 | 0% | 1,269 | 2,414 | +90% | 0 | 0 | — |
case-20 | fail→fail | 11,070 | 48,189 | +335% | 1 | 1 | 0% | 2,297 | 4,332 | +89% | 0 | 0 | — |
case-21 | fail→pass | 13,457 | 6,630 | -51% | 1 | 1 | 0% | 2,097 | 2,760 | +32% | 0 | 0 | — |
case-22 | fail→fail | 27,324 | 6,148 | -77% | 1 | 1 | 0% | 911 | 2,926 | +221% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 21 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +59 percentage points is the difference between those two pass rates over the 21 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
| Model | Method | Date | Lift |
|---|---|---|---|
| gemini-3.6-flash | verified | 7/28/2026 | +50% |
Other measured skills in the registry, with their headline benchmark lift.