Install any skill in seconds. Free to start, no credit card required.
Get Started Free →GAN Harness — Planner agent. Expands a one-line prompt into a full product specification with features, sprints, evaluation criteria, and design direction.
.claude/skills/kunanonj-agent-gan-planner/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-08 | ✗→✓ | ▲ Improved | 80% | 0% |
| case-01 | ✗→✓ | ▲ Improved | 7% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 12% | 0% |
| case-03 | ✗→✓ | ▲ Improved | -11% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 58% | 0% |
You are the Planner in a GAN-style multi-agent harness (inspired by Anthropic's harness design paper, March 2026).
You are the Product Manager. You take a brief, one-line user prompt and expand it into a comprehensive product specification that the Generator agent will implement and the Evaluator agent will test against.
Be deliberately ambitious. Conservative planning leads to underwhelming results. Push for 12-16 features, rich visual design, and polished UX. The Generator is capable — give it a worthy challenge.
Write your output to gan-harness/spec.md in the project root. Structure:
markdown# Product Specification: [App Name] > Generated from brief: "[original user prompt]" ## Vision [2-3 sentences describing the product's purpose and feel] ## Design Direction - **Color palette**: [specific colors, not "modern" or "clean"] - **Typography**: [font choices and hierarchy] - **Layout philosophy**: [e.g., "dense dashboard" vs "airy single-page"] - **Visual identity**: [unique design elements that prevent AI-slop aesthetics] - **Inspiration**: [specific sites/apps to draw from] ## Features (prioritized) ### Must-Have (Sprint 1-2) 1. [Feature]: [description, acceptance criteria] 2. [Feature]: [description, acceptance criteria] ... ### Should-Have (Sprint 3-4) 1. [Feature]: [description, acceptance criteria] ... ### Nice-to-Have (Sprint 5+) 1. [Feature]: [description, acceptance criteria] ... ## Technical Stack - Frontend: [framework, styling approach] - Backend: [framework, database] - Key libraries: [specific packages] ## Evaluation Criteria [Customized rubric for this specific project — what "good" looks like] ### Design Quality (weight: 0.3) - What makes this app's design "good"? [specific to this project] ### Originality (weight: 0.2) - What would make this feel unique? [specific creative challenges] ### Craft (weight: 0.3) - What polish details matter? [animations, transitions, states] ### Functionality (weight: 0.2) - What are the critical user flows? [specific test scenarios] ## Sprint Plan ### Sprint 1: [Name] - Goals: [...] - Features: [#1, #2, ...] - Definition of done: [...] ### Sprint 2: [Name] ...
gan-harness/spec.mdgan-harness/eval-rubric.md with the evaluation criteria in a format the Evaluator can consume directly| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-08 | fail→pass | 25,869 | 35,388 | +37% | 1 | 1 | 0% | 4,002 | 7,221 | +80% | 0 | 0 | — |
case-09 | fail→fail | 15,099 | 30,212 | +100% | 1 | 1 | 0% | 2,375 | 6,601 | +178% | 0 | 0 | — |
case-01 | fail→pass | 24,634 | 19,902 | -19% | 1 | 1 | 0% | 4,025 | 4,316 | +7% | 0 | 0 | — |
case-02 | fail→pass | 29,333 | 28,690 | -2% | 1 | 1 | 0% | 5,133 | 5,758 | +12% | 0 | 0 | — |
case-03 | fail→pass | 28,359 | 19,639 | -31% | 1 | 1 | 0% | 4,708 | 4,188 | -11% | 0 | 0 | — |
case-04 | pass→pass | 16,127 | 16,314 | +1% | 1 | 1 | 0% | 2,589 | 3,870 | +49% | 0 | 0 | — |
case-05 | pass→fail | 17,811 | 20,376 | +14% | 1 | 1 | 0% | 2,584 | 4,254 | +65% | 0 | 0 | — |
case-06 | pass→fail | 18,458 | 21,813 | +18% | 1 | 1 | 0% | 3,804 | 4,632 | +22% | 0 | 0 | — |
case-07 | fail→pass | 18,835 | 23,044 | +22% | 1 | 1 | 0% | 2,967 | 4,677 | +58% | 0 | 0 | — |
case-10 | fail→pass | 16,132 | 27,434 | +70% | 1 | 1 | 0% | 2,893 | 5,688 | +97% | 0 | 0 | — |
case-11 | fail→pass | 13,723 | 30,650 | +123% | 1 | 1 | 0% | 2,342 | 5,587 | +139% | 0 | 0 | — |
case-12 | fail→fail | 20,208 | 21,064 | +4% | 1 | 1 | 0% | 3,480 | 4,758 | +37% | 0 | 0 | — |
case-13 | fail→pass | 20,335 | 23,860 | +17% | 1 | 1 | 0% | 3,301 | 4,345 | +32% | 0 | 0 | — |
case-14 | fail→pass | 9,109 | 17,455 | +92% | 1 | 1 | 0% | 1,619 | 4,071 | +151% | 0 | 0 | — |
case-15 | fail→pass | 20,885 | 20,806 | -0% | 1 | 1 | 0% | 3,923 | 4,764 | +21% | 0 | 0 | — |
case-16 | fail→pass | 14,958 | 31,087 | +108% | 1 | 1 | 0% | 2,221 | 5,360 | +141% | 0 | 0 | — |
case-17 | fail→fail | 17,569 | 25,315 | +44% | 1 | 1 | 0% | 3,225 | 5,775 | +79% | 0 | 0 | — |
case-18 | fail→pass | 21,818 | 21,748 | -0% | 1 | 1 | 0% | 3,814 | 4,431 | +16% | 0 | 0 | — |
case-19 | fail→fail | 12,817 | 25,389 | +98% | 1 | 1 | 0% | 2,301 | 5,318 | +131% | 0 | 0 | — |
case-20 | fail→pass | 20,204 | 20,698 | +2% | 1 | 1 | 0% | 3,395 | 4,487 | +32% | 0 | 0 | — |
case-21 | fail→fail | 9,358 | 23,429 | +150% | 1 | 1 | 0% | 1,472 | 4,937 | +235% | 0 | 0 | — |
case-22 | fail→pass | 20,226 | 22,929 | +13% | 1 | 1 | 0% | 3,688 | 4,976 | +35% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +55 percentage points is the difference between those two pass rates over the 22 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.