Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Run `@metaharness/darwin security bench` (upstream "Darwin Shield" / ADR-155) — evolves a champion security-detection harness against a 10-vuln / 9-decoy corpus and grades it on TPR/FPR/patch-pass/repro/unsafe vs four baselines (B0 static, B1 LLM-single-pass, B2 fixed-agent, B3 Darwin-champion). Closest reference implementation for ruflo's own ADR-155 nightly self-learning security harness (PR
.claude/skills/ruvnet-harness-security-bench/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-09 | ✗→✓ | ▲ Improved | 160% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -4% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 490% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 61% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 152% | 0% |
Surfaces the upstream metaharness-darwin security bench command. This is the upstream's own ADR-155 — Darwin Shield — and is the closest reference implementation for ruflo's nightly self-learning security harness (#2417).
ruflo's ADR-155 proposes three learning loops (per-dimension confidence, severity calibration, auto-fix bid). Loop A trains on accumulated (finding, dimension, human_outcome) tuples — but the gradient signal is only sound if the underlying detection mechanism converges on a known-good corpus. Darwin Shield evolves exactly that mechanism on a 10-vuln/9-decoy ground-truth set. Running this nightly gives us:
TPR=1/FPR=0 on the bench corpus, our Loop A's reward signal is noise.
when the security landscape (or our mutator policy) shifts.
points to weight per-dimension confidence against.
Implementation: scripts/security-bench.mjs.
npx -y @metaharness/darwin@~0.8.0 metaharness-darwin security bench --population N --cycles N [--seed S].3s × 19 evaluations × population × cycles + 30s overhead.At default --population 2 --cycles 1 ≈ 144s; at --population 4 --cycles 3 ≈ 12 min.
pass/fail rows (gate examples: "TPR improvement ≥ 25% vs fixed", "FPR reduction ≥ 40%", "Patch-test pass rate ≥ 80%", "Reproduction success ≥ 90%", "Unsafe outputs = 0", "Cost increase ≤ 2× fixed", "Beyond SOTA: champion statistically beats previous champion", "Compounding: false-positive repeat-rate drop ≥ 35%").
repro/unsafe/cost per harness).
--alert-on-fail, exit 1 when overall = FAIL.json{ "success": true, "data": { "overall": { "ok": true, "icon": "✅" }, "gates": { "total": 11, "passed": 11, "failed": 0, "details": [{ "ok": true, "criterion": "TPR improvement ≥ 25% vs fixed harness", "measured": "+150% (B2 0.4 → B3 1)" }, ...] }, "baselines": [ { "harness": "static-only", "fitness": 0.5665, "tpr": 0.3, "fpr": 1, "unsafe": 0, ... }, { "harness": "LLM single-pass", "fitness": 0.1365, ... }, { "harness": "fixed agent", "fitness": 0.598, ... }, { "harness": "Darwin champion", "fitness": 0.93275, "tpr": 1, "fpr": 0, ... } ], "rawMarkdown": "...", "shape": { "population": 2, "cycles": 1, "seed": null }, "durationMs": 142000 } }
The ADR-155 nightly workflow (per #2418 task W1.5) will spawn this as one of the active-pentest dimension's calls — its results become a trajectory record:
jsonc{ "dimension": "mcp-pentest", "subdimension": "darwin-shield-bench", "champion_fitness": 0.93275, "champion_tpr": 1, "champion_fpr": 0, "gates_passed": 11, "gates_failed": 0, "shape": { "population": 4, "cycles": 3 } }
Loop A learns: if darwin-shield-bench consistently passes on the seeded corpus, weight findings caught only by mcp-pentest higher.
| Code | Meaning | |---|---| | 0 | Bench ran (overall PASS or FAIL — distinguish via JSON overall.ok), or degraded | | 1 | --alert-on-fail and overall.ok === false | | 2 | Config error or upstream infrastructure failure |
When @metaharness/darwin is absent, emits {degraded: true, reason: 'metaharness-darwin-not-available'} and exits 0.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-09 | fail→pass | 16,482 | 10,810 | -34% | 1 | 1 | 0% | 1,170 | 3,037 | +160% | 0 | 0 | — |
case-01 | fail→fail | 10,791 | 15,067 | +40% | 1 | 1 | 0% | 2,037 | 1,482 | -27% | 0 | 0 | — |
case-02 | fail→fail | 15,606 | 25,144 | +61% | 1 | 1 | 0% | 3,646 | 6,555 | +80% | 0 | 0 | — |
case-03 | fail→fail | 11,417 | 12,025 | +5% | 1 | 1 | 0% | 2,961 | 1,460 | -51% | 0 | 0 | — |
case-04 | pass→pass | 10,328 | 3,355 | -68% | 1 | 1 | 0% | 1,772 | 1,752 | -1% | 0 | 0 | — |
case-05 | pass→pass | 6,049 | 2,378 | -61% | 1 | 1 | 0% | 1,020 | 1,685 | +65% | 0 | 0 | — |
case-06 | pass→pass | 14,295 | 3,138 | -78% | 1 | 1 | 0% | 2,510 | 1,420 | -43% | 0 | 0 | — |
case-07 | fail→pass | 10,166 | 2,662 | -74% | 1 | 1 | 0% | 1,708 | 1,646 | -4% | 0 | 0 | — |
case-08 | fail→fail | 9,283 | 4,046 | -56% | 1 | 1 | 0% | 1,754 | 1,786 | +2% | 0 | 0 | — |
case-10 | fail→pass | 3,756 | 3,377 | -10% | 1 | 1 | 0% | 293 | 1,728 | +490% | 0 | 0 | — |
case-11 | fail→pass | 16,445 | 2,786 | -83% | 1 | 1 | 0% | 948 | 1,528 | +61% | 0 | 0 | — |
case-12 | pass→pass | 6,247 | 1,855 | -70% | 1 | 1 | 0% | 1,089 | 1,486 | +36% | 0 | 0 | — |
case-13 | fail→pass | 41,212 | 1,356 | -97% | 1 | 1 | 0% | 573 | 1,445 | +152% | 0 | 0 | — |
case-14 | fail→pass | 9,024 | 3,115 | -65% | 1 | 1 | 0% | 1,437 | 1,500 | +4% | 0 | 0 | — |
case-15 | fail→pass | 16,286 | 1,248 | -92% | 1 | 1 | 0% | 900 | 1,441 | +60% | 0 | 0 | — |
case-16 | fail→pass | 8,235 | 1,726 | -79% | 1 | 1 | 0% | 1,420 | 1,555 | +10% | 0 | 0 | — |
case-17 | fail→pass | 16,519 | 4,377 | -74% | 1 | 1 | 0% | 2,874 | 1,782 | -38% | 0 | 0 | — |
case-18 | pass→pass | 8,451 | 3,125 | -63% | 1 | 1 | 0% | 1,472 | 1,429 | -3% | 0 | 0 | — |
case-19 | fail→pass | 9,594 | 1,215 | -87% | 1 | 1 | 0% | 1,726 | 1,447 | -16% | 0 | 0 | — |
case-20 | pass→pass | 14,134 | 13,181 | -7% | 1 | 1 | 0% | 2,581 | 3,995 | +55% | 0 | 0 | — |
case-21 | pass→pass | 8,414 | 9,017 | +7% | 1 | 1 | 0% | 1,723 | 3,030 | +76% | 0 | 0 | — |
case-22 | fail→fail | 16,844 | 15,000 | -11% | 1 | 1 | 0% | 3,359 | 4,329 | +29% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 17 counted toward the lift figure. The other 5 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +45 percentage points is the difference between those two pass rates over the 17 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.