Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Manage `@metaharness/darwin` bench suites — `bench create <repo>` scaffolds a JSON suite from a repo's test corpus; `bench verify <suite.json>` checks suite well-formedness. Bench suites are the fixed evaluation corpora that `harness-evolve --bench <suite.json>` scores variants against, decoupling evolution from the repo's natural tests. Degrades gracefully when @metaharness/darwin is absent.
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-09 | ✗→✓ | ▲ Improved | 1% | 0% |
| case-10 | ✗→✓ | ▲ Improved | -43% | 0% |
| case-11 | ✗→✓ | ▲ Improved | -62% | 0% |
| case-15 | ✗→✓ | ▲ Improved | -50% | 0% |
| case-16 | ✗→✓ | ▲ Improved | -30% | 0% |
Surfaces metaharness-darwin bench <create|verify> — the supporting verb for harness-evolve --bench. Use when you want evolution scored against a fixed corpus (independent of npm test) so champion fitness is comparable across commits or across forks of the same harness.
npm test isflaky, slow, or undersized — scaffold a deterministic bench suite once, then evolve against it repeatedly.
bench verify the checked-in suite on every PR that touches it(cheap; ~5s).
the evaluation without losing comparability to the parent.
Implementation: scripts/bench.mjs.
--op create--repo path; reject if missing.metaharness-darwin bench create <repo> [--out <suite.json>].<repo>/.metaharness/bench/suite.json (chosen by upstream).{ input, expectedOutput, weight } tasksderived from existing test cases.
--op verify--suite path; reject if missing.metaharness-darwin bench verify <suite.json>.json{ "success": true, "data": { "op": "verify", "taskCount": 42, "wellFormed": true, "durationMs": 870 } }
| Code | Meaning | |---|---| | 0 | OK (or degraded — Darwin absent) | | 1 | --op verify and suite malformed | | 2 | Config error or upstream invocation failure |
When @metaharness/darwin is absent, emits the standard {degraded: true, reason: 'metaharness-darwin-not-available'} payload and exits 0.
Other measured skills in the registry, with their headline benchmark lift.