Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Run a GEPA learning cycle via `metaharness learn` (upstream ADR-235, metaharness@0.3.0) — optimizes a harness genome against a SWE-bench-style slice manifest. $0 dry-run by default; `--run` is the explicit spend opt-in. Requires a metaharness repo checkout (`--repo` or $METAHARNESS_REPO) — without one it reports `checkout-required` with clone instructions. Degrades gracefully when metaharness is absent.
.claude/skills/ruvnet-harness-learn/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | -45% | 0% |
| case-04 | ✗→✓ | ▲ Improved | -2% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 16% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -16% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 13% | 0% |
Surfaces metaharness learn — the upstream GEPA learning harness that evolves harness policy genomes against a scored task corpus instead of hand-editing prompts. Candidates are scored on held-out slices and only measured winners promote (the shipped cand-6 genome is the first such promotion: holdout gold 2/12 → 3/12, zero regressions).
measured improvement loop rather than manual prompt iteration.
resolves the slice manifest and reports cost without any model calls.
harness-gepa --op renderto inspect what the promoted policy actually says.
The learning harness (GEPA + SWE-bench + Docker) is too heavy for the npm package, so learn needs a local clone:
bashgit clone https://github.com/ruvnet/metaharness.git node scripts/learn.mjs --repo ./metaharness --host claude-code --model haiku --slice slices/lite.json
Without a checkout the script emits {status: "checkout-required"} and exits 0 — a precondition report, not an error (distinct from degraded: true, which means the npm package itself is absent). The managed-service path (gateway-side learn jobs, no checkout) is upstream's ADR-235 follow-up and not available yet.
Implementation: scripts/learn.mjs.
--repo exists when given; export it as $METAHARNESS_REPO.metaharness binary (metaharness@~0.3.0, local installor one-time versioned cache — never @latest): metaharness learn --host <h> --model <m> --slice <s> [--run] via _harness.mjs (graceful degradation, hard timeout).
--run — real runs on largerslices need an explicit --timeout-ms matched to slice size × model cost.
the raw report text under rawReport.
--run is the ONLY path that spends. Everything else — dry-run, checkout probe, degraded path — is $0. The MCP tool (metaharness_learn) has a 120s subprocess budget; run real learning cycles from a terminal via ruflo metaharness learn ... --run --timeout-ms <big>.
0 — report produced (or dry-run, checkout-required, degraded)1 — --alert-on-fail and the learn run reported failure2 — config error (bad --repo path)| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 9,137 | 4,828 | -47% | 1 | 1 | 0% | 1,994 | 1,103 | -45% | 0 | 0 | — |
case-02 | fail→fail | 5,628 | 6,016 | +7% | 1 | 1 | 0% | 955 | 1,000 | +5% | 0 | 0 | — |
case-03 | fail→fail | 6,579 | 6,199 | -6% | 1 | 1 | 0% | 1,350 | 1,471 | +9% | 0 | 0 | — |
case-04 | fail→pass | 6,935 | 2,731 | -61% | 1 | 1 | 0% | 1,297 | 1,277 | -2% | 0 | 0 | — |
case-05 | pass→pass | 10,686 | 8,364 | -22% | 1 | 1 | 0% | 2,163 | 2,340 | +8% | 0 | 0 | — |
case-06 | fail→pass | 4,440 | 1,082 | -76% | 1 | 1 | 0% | 795 | 919 | +16% | 0 | 0 | — |
case-07 | fail→pass | 9,209 | 3,577 | -61% | 1 | 1 | 0% | 1,670 | 1,405 | -16% | 0 | 0 | — |
case-08 | fail→pass | 7,657 | 4,043 | -47% | 1 | 1 | 0% | 1,406 | 1,582 | +13% | 0 | 0 | — |
case-09 | fail→pass | 7,787 | 2,764 | -65% | 1 | 1 | 0% | 1,362 | 1,286 | -6% | 0 | 0 | — |
case-10 | pass→pass | 9,321 | 1,243 | -87% | 1 | 1 | 0% | 1,725 | 951 | -45% | 0 | 0 | — |
case-11 | fail→pass | 13,494 | 1,771 | -87% | 1 | 1 | 0% | 2,621 | 1,078 | -59% | 0 | 0 | — |
case-12 | fail→pass | 12,697 | 1,194 | -91% | 1 | 1 | 0% | 2,664 | 905 | -66% | 0 | 0 | — |
case-13 | fail→pass | 12,045 | 4,598 | -62% | 1 | 1 | 0% | 2,195 | 1,653 | -25% | 0 | 0 | — |
case-14 | pass→pass | 3,928 | 1,894 | -52% | 1 | 1 | 0% | 746 | 1,066 | +43% | 0 | 0 | — |
case-15 | fail→pass | 9,522 | 3,344 | -65% | 1 | 1 | 0% | 1,855 | 1,350 | -27% | 0 | 0 | — |
case-16 | fail→pass | 10,689 | 1,524 | -86% | 1 | 1 | 0% | 2,331 | 1,021 | -56% | 0 | 0 | — |
case-17 | fail→pass | 16,083 | 1,987 | -88% | 1 | 1 | 0% | 1,434 | 1,030 | -28% | 0 | 0 | — |
case-18 | pass→pass | 5,099 | 1,891 | -63% | 1 | 1 | 0% | 1,022 | 1,021 | -0% | 0 | 0 | — |
case-19 | fail→pass | 11,921 | 3,577 | -70% | 1 | 1 | 0% | 1,901 | 1,346 | -29% | 0 | 0 | — |
case-20 | fail→pass | 5,221 | 2,280 | -56% | 1 | 1 | 0% | 1,031 | 1,147 | +11% | 0 | 0 | — |
case-21 | fail→pass | 11,326 | 1,734 | -85% | 1 | 1 | 0% | 2,208 | 1,056 | -52% | 0 | 0 | — |
case-22 | fail→pass | 9,828 | 4,879 | -50% | 1 | 1 | 0% | 1,749 | 1,752 | +0% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 19 counted toward the lift figure. The other 3 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +73 percentage points is the difference between those two pass rates over the 19 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.