Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Inspect and audit GEPA genomes via the `@metaharness/darwin/gepa` library entry (darwin 0.8.0) — load/validate a genome (default: the shipped cand-6 promotion), render the system prompt a genome compiles to, or classify failure modes in a run transcript. The `gepaOptimize` loop itself is library-only (bring your own evaluator) and not surfaced here — use `harness-evolve` for sandbox-scored evolution. Degrades gracefully when @metaharness/darwin is absent.
.claude/skills/ruvnet-harness-gepa/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-19 | ✗→✓ | ▲ Improved | 4% | 0% |
| case-04 | ✗→✓ | ▲ Improved | -21% | 0% |
| case-05 | ✗→✓ | ▲ Improved | -43% | 0% |
| case-06 | ✗→✓ | ▲ Improved | -53% | 0% |
| case-09 | ✗→✓ | ▲ Improved | -30% | 0% |
Surfaces the GEPA (genetic-evolution prompt-adaptation) library exports from @metaharness/darwin/gepa. Unlike the other skills in this plugin there is no CLI binary behind this — the script dynamic-imports the library (local resolution first, versioned cache install as fallback) and calls the subprocess-safe subset.
--op render shows the actual systemprompt a genome compiles to — read THAT, not the raw JSON, before wiring a genome into a harness.
--op genome loads + validates the shippedcand-6 genome (first holdout-confirmed cheap-tier promotion; provenance ships in the package) or any genome file you point at.
--op validate --alert-on-invalid exits 1on structural errors.
--op analyze --transcript run.json classifiesfailure modes (GEPA's failure-class taxonomy) from a transcript array.
gepaOptimize — the optimization loop takes an in-process evaluate(candidate) callback ("bring your own evaluator") that cannot cross a subprocess boundary. Two supported paths instead:
import { gepaOptimize, loadCand6Genome } from '@metaharness/darwin/gepa'harness-evolve (darwin CLI evolve),which pairs GEPA with its own sandbox evaluators.
Implementation: scripts/gepa.mjs.
import('@metaharness/darwin/gepa'); on MODULE_NOT_FOUND fall back to aone-time npm install --prefix ~/.ruflo/darwin-cache-0.8.0 and import the cached dist/gepa/index.js (versioned dir → pin bumps invalidate).
--op:genome → loadGenome(fs, path) or loadCand6Genome() + validateGenomevalidate → validateGenome(rawJson) (raw parse so broken files reachthe validator instead of throwing in the loader)
render → buildSystemFromGenome(genome, ext?, glob?)analyze → analyzeTranscript(entries)--alert-on-invalid, 2 on bad input).bashnode scripts/gepa.mjs --op genome # cand-6 + validation node scripts/gepa.mjs --op render | jq -r .system # what does cand-6 SAY? node scripts/gepa.mjs --op validate --path my-genome.json --alert-on-invalid node scripts/gepa.mjs --op analyze --transcript run.json
0 — op completed (or degraded — darwin not installable)1 — --alert-on-invalid and validation found errors2 — config error (unknown op, missing/broken input file)| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 8,211 | 4,895 | -40% | 1 | 1 | 0% | 1,270 | 956 | -25% | 0 | 0 | — |
case-19 | fail→pass | 8,018 | 3,489 | -56% | 1 | 1 | 0% | 1,387 | 1,445 | +4% | 0 | 0 | — |
case-02 | fail→fail | 7,955 | 4,559 | -43% | 1 | 1 | 0% | 1,436 | 976 | -32% | 0 | 0 | — |
case-03 | fail→fail | 2,990 | 4,715 | +58% | 1 | 1 | 0% | 475 | 925 | +95% | 0 | 0 | — |
case-04 | fail→pass | 17,614 | 2,966 | -83% | 1 | 1 | 0% | 1,652 | 1,298 | -21% | 0 | 0 | — |
case-05 | fail→pass | 17,688 | 5,901 | -67% | 1 | 1 | 0% | 3,406 | 1,947 | -43% | 0 | 0 | — |
case-06 | fail→pass | 14,806 | 3,262 | -78% | 1 | 1 | 0% | 2,841 | 1,347 | -53% | 0 | 0 | — |
case-07 | pass→pass | 18,132 | 2,029 | -89% | 1 | 1 | 0% | 1,617 | 1,041 | -36% | 0 | 0 | — |
case-08 | pass→pass | 5,650 | 1,363 | -76% | 1 | 1 | 0% | 888 | 917 | +3% | 0 | 0 | — |
case-09 | fail→pass | 10,975 | 3,670 | -67% | 1 | 1 | 0% | 1,958 | 1,380 | -30% | 0 | 0 | — |
case-10 | fail→pass | 13,972 | 3,079 | -78% | 1 | 1 | 0% | 2,332 | 1,256 | -46% | 0 | 0 | — |
case-11 | pass→pass | 10,994 | 4,084 | -63% | 1 | 1 | 0% | 1,738 | 1,429 | -18% | 0 | 0 | — |
case-12 | fail→pass | 8,482 | 2,767 | -67% | 1 | 1 | 0% | 1,361 | 1,244 | -9% | 0 | 0 | — |
case-13 | fail→pass | 13,586 | 5,446 | -60% | 1 | 1 | 0% | 2,296 | 1,702 | -26% | 0 | 0 | — |
case-14 | fail→pass | 8,154 | 1,859 | -77% | 1 | 1 | 0% | 1,376 | 1,083 | -21% | 0 | 0 | — |
case-15 | fail→pass | 11,332 | 2,269 | -80% | 1 | 1 | 0% | 1,940 | 1,151 | -41% | 0 | 0 | — |
case-16 | fail→pass | 19,018 | 2,032 | -89% | 1 | 1 | 0% | 1,321 | 1,098 | -17% | 0 | 0 | — |
case-17 | fail→pass | 14,354 | 1,580 | -89% | 1 | 1 | 0% | 2,590 | 987 | -62% | 0 | 0 | — |
case-18 | pass→pass | 26,595 | 3,628 | -86% | 1 | 1 | 0% | 2,364 | 1,414 | -40% | 0 | 0 | — |
case-20 | pass→pass | 5,081 | 5,371 | +6% | 1 | 1 | 0% | 876 | 1,784 | +104% | 0 | 0 | — |
case-21 | pass→fail | 4,027 | 4,805 | +19% | 1 | 1 | 0% | 767 | 1,493 | +95% | 0 | 0 | — |
case-22 | pass→pass | 17,531 | 11,307 | -36% | 1 | 1 | 0% | 2,246 | 2,757 | +23% | 0 | 0 | — |
case-23 | pass→pass | 4,378 | 3,303 | -25% | 1 | 1 | 0% | 752 | 1,233 | +64% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 20 counted toward the lift figure. The other 3 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +48 percentage points is the difference between those two pass rates over the 20 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.