Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Run `@metaharness/darwin evolve <repo>` to mutate a harness's seven policy surfaces (planner/contextBuilder/reviewer/retryPolicy/toolPolicy/memoryPolicy/scorePolicy), sandbox-score each variant, and promote only measured wins. The model is frozen; the harness evolves. Closes the loop ADR-150 opens (score+genome describe; evolve changes). Degrades gracefully when @metaharness/darwin is absent (ADR-150 + ADR-153 architectural constraints).
.claude/skills/ruvnet-harness-evolve/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-06 | ✗→✓ | ▲ Improved | 53% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 22% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 35% | 0% |
| case-10 | ✗→✓ | ▲ Improved | -27% | 0% |
| case-11 | ✗→✓ | ▲ Improved | -25% | 0% |
Surfaces the upstream metaharness-darwin evolve CLI as a ruflo skill. The write layer that pairs with ADR-150's read layer (score / genome / mcp-scan / threat-model / oia-audit). Use when you have a harness whose readiness scores are flat and you want to discover which surface mutation moves them — without retraining the foundation model.
harness-score result is below target and you don't know which policysurface is responsible.
starting configuration empirically rather than hand-tuning.
(treat darwin's champion as the strawman).
Wire it into CI for one-shot exploration, not for autonomous self-modification.
— the CI gate verifies graceful degradation, not convergence.
Implementation: scripts/evolve.mjs.
--repo exists, caps on --generations ≤ 50, --children≤ 20, --concurrency ≤ 8, sandbox/selection/mutator are known values).
--confirm: print plan + exit 0 (mirrors harness-mint safetyconvention; defense in depth over the upstream safety.ts checks).
--confirm: shell to npx -y @metaharness/darwin@~0.8.0 metaharness-darwin evolve <repo> ...via the shared _darwin.mjs async helper. Per-generation progress is forwarded to stderr; final champion JSON is captured from stdout.
generations × children × per-variant (per-variant≈ 60s real, ≈ 2s mock). Caller may override with --timeout-ms.
remap. This is a designed-in tripwire (a variant tripped inspectVariant for secrets / shell-out / network / dynamic-eval). See ADR-153 §"Safety model".
--alert-on-no-improvement: exit 1 when champion ≤ parent.| Surface | What it owns | |---|---| | planner | task decomposition / step ordering | | contextBuilder | what gets fed into the prompt | | reviewer | self-critique / output verification | | retryPolicy | when + how to retry on failure | | toolPolicy | which tools the agent may use, under which conditions | | memoryPolicy | what to persist, recall, forget | | scorePolicy | how the agent grades its own output |
One mutation per variant. Multi-surface mutations are not allowed (causal attribution stays clean).
Reports land under <repo>/.metaharness/:
.metaharness/
archive.json # full lineage tree (sampling next gen draws from this)
lineage.json # parent→child edges only
variants/<id>/ # per-variant code (kept for audit)
runs/<id>/ # per-variant sandbox test output
reports/winner.json # final champion + score delta vs parentSkill stdout = JSON {success, data: {champion, plan, durationMs, improved}} (plus data.diagnosis when --diagnose is passed — see below).
--diagnose)GEPA's key trick is natural-language failure diagnosis from execution traces feeding the next mutation — not just scalar fitness. --diagnose adds a modest slice of that: after the evolution completes, the losing / failed variants' transcripts are run through darwin's GEPA library ops (analyzeTranscript + classifyFailure, via the shared importGepa resolver in scripts/_darwin.mjs) and a diagnosis section is appended to the emitted JSON:
json"diagnosis": { "available": true, "scope": "losing-variants", "variants": [ { "id": "g1_v0", "transcripts": 2, "failureClasses": { "exploration-loop": 1, "edit-mechanics": 1 }, "dominantClass": "exploration-loop" } ], "totals": { "exploration-loop": 1, "edit-mechanics": 1 } }
Upstream shape caveats (verified against @metaharness/darwin@0.8.0):
metaharness-darwin evolve --json prints a TEXT leaderboard — the stdoutcarries no JSON and no transcripts. Per-variant run records live at <repo>/.metaharness/runs/<id>.json.
stderr}), which are NOT GEPA {actionRaw, obs} transcripts. Diagnosis therefore uses GEPA-shaped transcripts when a run record embeds them (agent sandbox / future upstream), falls back to the champion's transcript, and otherwise emits diagnosis: {available: false, reason, traceSummary} where traceSummary is a mechanical per-variant tally (tasks / failed / timedOut / blockedActions).
--diagnose NEVER fails the run — any internal error degrades to{available: false, reason: "diagnosis-failed: ..."}.
| Code | Meaning | |---|---| | 0 | Evolved OK, or dry-run, or degraded (Darwin absent) | | 1 | --alert-on-no-improvement and champion did not beat parent | | 2 | Config error or evolution infrastructure failure | | 99 | Upstream "safety-disqualified" (PROPAGATED, not remapped) |
When @metaharness/darwin is not installed, the script emits {degraded: true, reason: 'metaharness-darwin-not-available', hint: ...} and exits 0. ruflo continues to function. CI's no-metaharness-smoke.yml-style job asserts this path.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-08 | pass→pass | 10,766 | 4,045 | -62% | 1 | 1 | 0% | 2,100 | 2,309 | +10% | 0 | 0 | — |
case-02 | fail→fail | 3,789 | 4,188 | +11% | 1 | 1 | 0% | 195 | 1,815 | +831% | 0 | 0 | — |
case-01 | fail→fail | 3,092 | 6,053 | +96% | 1 | 1 | 0% | 172 | 2,138 | +1143% | 0 | 0 | — |
case-03 | fail→fail | 6,793 | 4,583 | -33% | 1 | 1 | 0% | 1,819 | 1,848 | +2% | 0 | 0 | — |
case-04 | fail→fail | 4,926 | 4,600 | -7% | 1 | 1 | 0% | 936 | 1,900 | +103% | 0 | 0 | — |
case-05 | pass→fail | 7,264 | 4,519 | -38% | 1 | 1 | 0% | 1,306 | 1,889 | +45% | 0 | 0 | — |
case-06 | fail→pass | 6,928 | 2,858 | -59% | 1 | 1 | 0% | 1,432 | 2,196 | +53% | 0 | 0 | — |
case-07 | fail→pass | 15,023 | 1,747 | -88% | 1 | 1 | 0% | 1,520 | 1,860 | +22% | 0 | 0 | — |
case-09 | fail→pass | 10,019 | 4,676 | -53% | 1 | 1 | 0% | 1,740 | 2,345 | +35% | 0 | 0 | — |
case-10 | fail→pass | 11,401 | 1,965 | -83% | 1 | 1 | 0% | 2,614 | 1,908 | -27% | 0 | 0 | — |
case-11 | fail→pass | 14,436 | 4,722 | -67% | 1 | 1 | 0% | 2,595 | 1,937 | -25% | 0 | 0 | — |
case-12 | fail→pass | 9,347 | 4,080 | -56% | 1 | 1 | 0% | 1,874 | 2,487 | +33% | 0 | 0 | — |
case-13 | fail→pass | 9,185 | 2,359 | -74% | 1 | 1 | 0% | 1,565 | 1,956 | +25% | 0 | 0 | — |
case-14 | fail→pass | 10,951 | 2,098 | -81% | 1 | 1 | 0% | 2,215 | 1,951 | -12% | 0 | 0 | — |
case-15 | fail→pass | 7,489 | 2,473 | -67% | 1 | 1 | 0% | 1,423 | 1,971 | +39% | 0 | 0 | — |
case-16 | fail→pass | 7,340 | 3,265 | -56% | 1 | 1 | 0% | 1,326 | 2,198 | +66% | 0 | 0 | — |
case-17 | fail→pass | 17,759 | 1,893 | -89% | 1 | 1 | 0% | 1,637 | 1,867 | +14% | 0 | 0 | — |
case-18 | fail→pass | 7,013 | 6,209 | -11% | 1 | 1 | 0% | 1,437 | 2,722 | +89% | 0 | 0 | — |
case-19 | pass→pass | 6,541 | 4,357 | -33% | 1 | 1 | 0% | 1,227 | 2,288 | +86% | 0 | 0 | — |
case-20 | fail→pass | 14,418 | 7,062 | -51% | 1 | 1 | 0% | 2,940 | 2,847 | -3% | 0 | 0 | — |
case-21 | fail→pass | 2,457 | 892 | -64% | 1 | 1 | 0% | 318 | 1,621 | +410% | 0 | 0 | — |
case-22 | fail→pass | 11,321 | 3,789 | -67% | 1 | 1 | 0% | 2,120 | 2,230 | +5% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 15 counted toward the lift figure. The other 7 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +64 percentage points is the difference between those two pass rates over the 15 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.