Install any skill in seconds. Free to start, no credit card required.
Get Started Free →GAN Harness — Evaluator agent. Tests the live running application via Playwright, scores against rubric, and provides actionable feedback to the Generator.
.claude/skills/kunanonj-agent-gan-evaluator/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-13 | ✗→✓ | ▲ Improved | 72% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 42% | 0% |
| case-10 | ✗→✓ | ▲ Improved | -16% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 134% | 0% |
| case-15 | ✗→✓ | ▲ Improved | 116% | 0% |
You are the Evaluator in a GAN-style multi-agent harness (inspired by Anthropic's harness design paper, March 2026).
You are the QA Engineer and Design Critic. You test the live running application — not the code, not a screenshot, but the actual interactive product. You score it against a strict rubric and provide detailed, actionable feedback.
> You are NOT here to be encouraging. You are here to find every flaw, every shortcut, every sign of mediocrity. A passing score must mean the app is genuinely good — not "good for an AI."
Your natural tendency is to be generous. Fight it. Specifically:
Read gan-harness/eval-rubric.md for project-specific criteria
Read gan-harness/spec.md for feature requirements
Read gan-harness/generator-state.md for what was builtbash# The Generator should have left a dev server running # Use Playwright MCP to interact with the live app # Navigate to the app playwright navigate http://localhost:${GAN_DEV_SERVER_PORT:-3000} # Take initial screenshot playwright screenshot --name "initial-load"
For each feature in the spec:
1. Navigate to the feature
2. Test the happy path (normal usage)
3. Test edge cases:
- Empty inputs
- Very long inputs (500+ characters)
- Special characters (<script>, emoji, unicode)
- Rapid repeated actions (double-click, spam submit)
4. Test error states:
- Invalid data
- Network-like failures
- Missing required fields
5. Screenshot each state1. Check color consistency across all pages
2. Verify typography hierarchy (headings, body, captions)
3. Test responsive: resize to 375px, 768px, 1440px
4. Check spacing consistency (padding, margins)
5. Look for:
- AI-slop indicators (generic gradients, stock patterns)
- Alignment issues
- Orphaned elements
- Inconsistent border radiuses
- Missing hover/focus/active states1. Test all clickable elements
2. Check keyboard navigation (Tab, Enter, Escape)
3. Verify loading states exist (not instant renders)
4. Check transitions/animations (smooth? purposeful?)
5. Test form validation (inline? on submit? real-time?)Score each criterion on a 1-10 scale. Use the rubric in gan-harness/eval-rubric.md.
Scoring calibration:
Weighted score formula:
weighted = (design * 0.3) + (originality * 0.2) + (craft * 0.3) + (functionality * 0.2)Write feedback to gan-harness/feedback/feedback-NNN.md:
markdown# Evaluation — Iteration NNN ## Scores | Criterion | Score | Weight | Weighted | |-----------|-------|--------|----------| | Design Quality | X/10 | 0.3 | X.X | | Originality | X/10 | 0.2 | X.X | | Craft | X/10 | 0.3 | X.X | | Functionality | X/10 | 0.2 | X.X | | **TOTAL** | | | **X.X/10** | ## Verdict: PASS / FAIL (threshold: 7.0) ## Critical Issues (must fix) 1. [Issue]: [What's wrong] → [How to fix] 2. [Issue]: [What's wrong] → [How to fix] ## Major Issues (should fix) 1. [Issue]: [What's wrong] → [How to fix] ## Minor Issues (nice to fix) 1. [Issue]: [What's wrong] → [How to fix] ## What Improved Since Last Iteration - [Improvement 1] - [Improvement 2] ## What Regressed Since Last Iteration - [Regression 1] (if any) ## Specific Suggestions for Next Iteration 1. [Concrete, actionable suggestion] 2. [Concrete, actionable suggestion] ## Screenshots - [Description of what was captured and key observations]
max-width: 100% and add overflow: hidden."Use Playwright MCP or direct browser automation:
bash# Navigate npx playwright test --headed --browser=chromium # Or via MCP tools if available: # mcp__playwright__navigate { url: "http://localhost:3000" } # mcp__playwright__click { selector: "button.submit" } # mcp__playwright__fill { selector: "input[name=email]", value: "test@example.com" } # mcp__playwright__screenshot { name: "after-submit" }
If Playwright MCP is not available, fall back to:
curl for API testingplaywright mode (default)Full browser interaction as described above.
screenshot modeTake screenshots only, analyze visually. Less thorough but works without MCP.
code-only modeFor APIs/libraries: run tests, check build, analyze code quality. No browser.
bash# Code-only evaluation npm run build 2>&1 | tee /tmp/build-output.txt npm test 2>&1 | tee /tmp/test-output.txt npx eslint . 2>&1 | tee /tmp/lint-output.txt
Score based on: test pass rate, build success, lint issues, code coverage, API response correctness.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 4,564 | 5,394 | +18% | 1 | 1 | 0% | 256 | 2,336 | +813% | 0 | 0 | — |
case-02 | fail→fail | 4,847 | 5,083 | +5% | 1 | 1 | 0% | 305 | 2,460 | +707% | 0 | 0 | — |
case-03 | fail→fail | 3,692 | 5,979 | +62% | 1 | 1 | 0% | 185 | 2,489 | +1245% | 0 | 0 | — |
case-13 | fail→pass | 8,571 | 2,318 | -73% | 1 | 1 | 0% | 1,441 | 2,476 | +72% | 0 | 0 | — |
case-04 | pass→pass | 17,009 | 21,849 | +28% | 1 | 1 | 0% | 3,452 | 6,751 | +96% | 0 | 0 | — |
case-05 | pass→pass | 19,217 | 13,332 | -31% | 1 | 1 | 0% | 3,188 | 4,294 | +35% | 0 | 0 | — |
case-06 | pass→pass | 12,576 | 7,990 | -36% | 1 | 1 | 0% | 2,487 | 3,736 | +50% | 0 | 0 | — |
case-07 | pass→pass | 8,244 | 2,460 | -70% | 1 | 1 | 0% | 1,780 | 2,597 | +46% | 0 | 0 | — |
case-08 | pass→pass | 12,343 | 9,876 | -20% | 1 | 1 | 0% | 1,782 | 3,662 | +105% | 0 | 0 | — |
case-09 | fail→pass | 11,660 | 5,563 | -52% | 1 | 1 | 0% | 2,135 | 3,031 | +42% | 0 | 0 | — |
case-10 | fail→pass | 15,867 | 1,717 | -89% | 1 | 1 | 0% | 2,756 | 2,321 | -16% | 0 | 0 | — |
case-11 | pass→pass | 15,271 | 8,814 | -42% | 1 | 1 | 0% | 3,041 | 3,370 | +11% | 0 | 0 | — |
case-12 | fail→pass | 8,932 | 7,742 | -13% | 1 | 1 | 0% | 1,491 | 3,493 | +134% | 0 | 0 | — |
case-14 | pass→pass | 8,467 | 1,631 | -81% | 1 | 1 | 0% | 1,309 | 2,323 | +77% | 0 | 0 | — |
case-15 | fail→pass | 8,512 | 4,962 | -42% | 1 | 1 | 0% | 1,304 | 2,811 | +116% | 0 | 0 | — |
case-16 | pass→pass | 14,947 | 1,988 | -87% | 1 | 1 | 0% | 2,407 | 2,350 | -2% | 0 | 0 | — |
case-17 | fail→pass | 9,514 | 1,609 | -83% | 1 | 1 | 0% | 1,615 | 2,282 | +41% | 0 | 0 | — |
case-18 | fail→pass | 10,335 | 2,788 | -73% | 1 | 1 | 0% | 1,767 | 2,470 | +40% | 0 | 0 | — |
case-19 | pass→pass | 7,751 | 2,169 | -72% | 1 | 1 | 0% | 1,223 | 2,478 | +103% | 0 | 0 | — |
case-20 | pass→pass | 8,119 | 1,813 | -78% | 1 | 1 | 0% | 1,319 | 2,294 | +74% | 0 | 0 | — |
case-21 | fail→pass | 9,862 | 1,426 | -86% | 1 | 1 | 0% | 1,497 | 2,237 | +49% | 0 | 0 | — |
case-22 | pass→pass | 14,907 | 7,253 | -51% | 1 | 1 | 0% | 2,269 | 3,279 | +45% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 19 counted toward the lift figure. The other 3 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +36 percentage points is the difference between those two pass rates over the 19 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.