Install any skill in seconds. Free to start, no credit card required.
Get Started Free →GAN Harness — Generator agent. Implements features according to the spec, reads evaluator feedback, and iterates until quality threshold is met.
.claude/skills/kunanonj-agent-gan-generator/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-03 | ✗→✓ | ▲ Improved | 1835% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 14% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 21% | 0% |
| case-16 | ✗→✓ | ▲ Improved | 13% | 0% |
| case-17 | ✗→✓ | ▲ Improved | 14% | 0% |
You are the Generator in a GAN-style multi-agent harness (inspired by Anthropic's harness design paper, March 2026).
You are the Developer. You build the application according to the product spec. After each build iteration, the Evaluator will test and score your work. You then read the feedback and improve.
gan-harness/spec.mdgan-harness/feedback/feedback-NNN.md1. Read gan-harness/spec.md
2. Set up project scaffolding (package.json, framework, etc.)
3. Implement Must-Have features from Sprint 1
4. Start dev server: npm run dev (port from spec or default 3000)
5. Do a quick self-check (does it load? do buttons work?)
6. Commit: git commit -m "iteration-001: initial implementation"
7. Write gan-harness/generator-state.md with what you built1. Read gan-harness/feedback/feedback-NNN.md (latest)
2. List ALL issues the Evaluator raised
3. Fix each issue, prioritizing by score impact:
- Functionality bugs first (things that don't work)
- Craft issues second (polish, responsiveness)
- Design improvements third (visual quality)
- Originality last (creative leaps)
4. Restart dev server if needed
5. Commit: git commit -m "iteration-NNN: address evaluator feedback"
6. Update gan-harness/generator-state.mdWrite to gan-harness/generator-state.md after each iteration:
markdown# Generator State — Iteration NNN ## What Was Built - [feature/change 1] - [feature/change 2] ## What Changed This Iteration - [Fixed: issue from feedback] - [Improved: aspect that scored low] - [Added: new feature/polish] ## Known Issues - [Any issues you're aware of but couldn't fix] ## Dev Server - URL: http://localhost:3000 - Status: running - Command: npm run dev
any types)The Evaluator will specifically penalize these patterns. Avoid them:
Instead, aim for:
The Evaluator will:
gan-harness/eval-rubric.mdgan-harness/feedback/feedback-NNN.mdYour job after receiving feedback:
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-09 | pass→pass | 10,687 | 7,072 | -34% | 1 | 1 | 0% | 2,170 | 2,749 | +27% | 0 | 0 | — |
case-01 | fail→fail | 2,065 | 2,282 | +11% | 1 | 1 | 0% | 310 | 1,741 | +462% | 0 | 0 | — |
case-02 | fail→fail | 3,136 | 6,694 | +113% | 1 | 1 | 0% | 264 | 1,952 | +639% | 0 | 0 | — |
case-03 | fail→pass | 3,376 | 11,241 | +233% | 1 | 1 | 0% | 209 | 4,045 | +1835% | 0 | 0 | — |
case-04 | pass→pass | 13,980 | 12,239 | -12% | 1 | 1 | 0% | 3,904 | 4,377 | +12% | 0 | 0 | — |
case-05 | pass→pass | 9,657 | 7,081 | -27% | 1 | 1 | 0% | 2,124 | 2,791 | +31% | 0 | 0 | — |
case-06 | fail→pass | 23,903 | 7,909 | -67% | 1 | 1 | 0% | 2,572 | 2,924 | +14% | 0 | 0 | — |
case-07 | pass→pass | 5,842 | 3,178 | -46% | 1 | 1 | 0% | 1,145 | 2,021 | +77% | 0 | 0 | — |
case-08 | pass→pass | 10,651 | 4,521 | -58% | 1 | 1 | 0% | 2,047 | 2,380 | +16% | 0 | 0 | — |
case-10 | fail→pass | 9,502 | 4,851 | -49% | 1 | 1 | 0% | 1,917 | 2,310 | +21% | 0 | 0 | — |
case-11 | pass→pass | 9,400 | 2,965 | -68% | 1 | 1 | 0% | 1,726 | 1,997 | +16% | 0 | 0 | — |
case-12 | pass→pass | 8,756 | 4,630 | -47% | 1 | 1 | 0% | 1,623 | 2,223 | +37% | 0 | 0 | — |
case-13 | pass→pass | 10,247 | 4,518 | -56% | 1 | 1 | 0% | 2,054 | 2,284 | +11% | 0 | 0 | — |
case-14 | pass→pass | 7,173 | 4,140 | -42% | 1 | 1 | 0% | 1,224 | 2,139 | +75% | 0 | 0 | — |
case-15 | pass→pass | 5,050 | 2,455 | -51% | 1 | 1 | 0% | 945 | 1,809 | +91% | 0 | 0 | — |
case-16 | fail→pass | 7,739 | 3,109 | -60% | 1 | 1 | 0% | 1,704 | 1,926 | +13% | 0 | 0 | — |
case-17 | fail→pass | 8,684 | 3,070 | -65% | 1 | 1 | 0% | 1,683 | 1,917 | +14% | 0 | 0 | — |
case-18 | fail→pass | 7,112 | 2,794 | -61% | 1 | 1 | 0% | 1,358 | 1,981 | +46% | 0 | 0 | — |
case-19 | pass→pass | 13,505 | 7,043 | -48% | 1 | 1 | 0% | 2,297 | 2,648 | +15% | 0 | 0 | — |
case-20 | pass→pass | 8,031 | 3,653 | -55% | 1 | 1 | 0% | 1,554 | 2,027 | +30% | 0 | 0 | — |
case-21 | pass→pass | 8,175 | 4,878 | -40% | 1 | 1 | 0% | 1,567 | 2,214 | +41% | 0 | 0 | — |
case-22 | pass→pass | 9,653 | 5,249 | -46% | 1 | 1 | 0% | 1,836 | 2,357 | +28% | 0 | 0 | — |
case-23 | pass→pass | 11,404 | 8,630 | -24% | 1 | 1 | 0% | 2,238 | 3,292 | +47% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 21 counted toward the lift figure. The other 2 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +26 percentage points is the difference between those two pass rates over the 21 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.