Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Sign, verify, and track fix-marker regressions over time using a deterministic Ed25519 witness manifest. Works in any project — clone the toolkit, run init, register fixes, regen on each release.
.claude/skills/ruvnet-witness/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 26% | 0% |
| case-02 | ✗→✓ | ▲ Improved | -14% | 0% |
| case-03 | ✗→✓ | ▲ Improved | -9% | 0% |
| case-04 | ✗→✓ | ▲ Improved | -50% | 0% |
| case-06 | ✗→✓ | ▲ Improved | -54% | 0% |
The witness toolkit lets you ship every release with a signed manifest that lists every documented fix in your codebase along with a sha256 + marker substring. Anyone with the same git commit can re-derive the public key and verify the signature without a committed private key.
A temporal history (JSONL) tracks how the fix population evolves across releases — so when a regression appears, you can pinpoint the commit that introduced it, not just "it's broken now."
This skill works two ways:
.github/workflows/v3-ci.yml job witness-verify).
plugins/ruflo-core/scripts/witness/into your repo, run init.mjs, register your fixes in witness-fixes.json, and call regen.mjs from your release pipeline.
bash# One-time bootstrap — creates verification.md.json, # verification-history.jsonl, and witness-fixes.json template node plugins/ruflo-core/scripts/witness/init.mjs --root . # Edit witness-fixes.json: add { id, desc, file, marker } per fix. # A "marker" is a distinctive substring that MUST appear in `file` # while the fix is present. If someone reverts the fix, the marker # disappears and `verify` reports it as `regressed`. # Regenerate the manifest (signing requires @noble/ed25519) npm i @noble/ed25519 node plugins/ruflo-core/scripts/witness/regen.mjs \ --manifest verification.md.json \ --history verification-history.jsonl \ --fixes witness-fixes.json # Verify markers are present in the live tree node plugins/ruflo-core/scripts/witness/verify.mjs \ --manifest verification.md.json # Or authenticate the manifest and check source markers in a clean clone. # Generated dist/ entries are explicitly reported as skipped. node plugins/ruflo-core/scripts/witness/verify.mjs \ --manifest verification.md.json --source-only
bash# Latest snapshot vs. previous node plugins/ruflo-core/scripts/witness/history.mjs \ --history verification-history.jsonl summary # For each currently-regressed fix, find the commit that introduced it node plugins/ruflo-core/scripts/witness/history.mjs \ --history verification-history.jsonl regressions # Status timeline for a specific fix node plugins/ruflo-core/scripts/witness/history.mjs \ --history verification-history.jsonl timeline --id F1 # Machine-readable for CI node plugins/ruflo-core/scripts/witness/history.mjs \ --history verification-history.jsonl summary --json
summary exits non-zero if any fix newly regressed since the last snapshot — drop it in CI as a soft pre-merge gate.
verification.md.json — always regenerate via regen.mjs,otherwise the signature breaks.
'function', 'import') — pick somethingunique enough that grep doesn't false-positive against unrelated code.
--history, you lose theability to bisect when a regression was introduced.
verification.md.json andverification-history.jsonl belong in the same commit; the JSONL is what lets future you verify the signed manifest is the latest in the line.
scripts/witness/lib.mjs — shared regenerate / history logic.scripts/witness/regen.mjs — CLI: sign + append history.scripts/witness/history.mjs — CLI: query the temporal log.scripts/witness/init.mjs — CLI: bootstrap into a fresh project.scripts/witness/verify.mjs — CLI: validate signature + markers.v3-ci.yml job witness-verify runs after the behavioral smoke tests and before publish. Failure modes:
| Failure | Cause | |---|---| | signatureValid: no | manifest hand-edited; re-run regen | | regressed: > 0 | a documented fix lost its marker since issuance | | missing: > 0 | a cited dist file no longer exists; rebuild or remove the entry | | scope: source-only | signature + source markers checked; generated entries intentionally skipped |
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 7,079 | 2,332 | -67% | 1 | 1 | 0% | 1,253 | 1,576 | +26% | 0 | 0 | — |
case-02 | fail→pass | 9,942 | 3,438 | -65% | 1 | 1 | 0% | 2,056 | 1,763 | -14% | 0 | 0 | — |
case-03 | fail→pass | 9,565 | 2,822 | -70% | 1 | 1 | 0% | 1,883 | 1,717 | -9% | 0 | 0 | — |
case-04 | fail→pass | 16,264 | 2,802 | -83% | 1 | 1 | 0% | 3,416 | 1,717 | -50% | 0 | 0 | — |
case-05 | pass→pass | 6,388 | 1,542 | -76% | 1 | 1 | 0% | 1,251 | 1,402 | +12% | 0 | 0 | — |
case-06 | fail→pass | 18,147 | 2,475 | -86% | 1 | 1 | 0% | 3,399 | 1,554 | -54% | 0 | 0 | — |
case-07 | fail→pass | 11,716 | 2,479 | -79% | 1 | 1 | 0% | 2,151 | 1,610 | -25% | 0 | 0 | — |
case-08 | fail→pass | 8,511 | 2,269 | -73% | 1 | 1 | 0% | 1,510 | 1,528 | +1% | 0 | 0 | — |
case-09 | pass→pass | 13,774 | 3,974 | -71% | 1 | 1 | 0% | 2,309 | 1,965 | -15% | 0 | 0 | — |
case-10 | fail→pass | 7,292 | 3,457 | -53% | 1 | 1 | 0% | 1,162 | 1,762 | +52% | 0 | 0 | — |
case-11 | fail→pass | 8,641 | 6,702 | -22% | 1 | 1 | 0% | 1,496 | 2,162 | +45% | 0 | 0 | — |
case-12 | fail→pass | 6,298 | 4,233 | -33% | 1 | 1 | 0% | 969 | 1,861 | +92% | 0 | 0 | — |
case-13 | fail→pass | 8,840 | 2,799 | -68% | 1 | 1 | 0% | 1,532 | 1,585 | +3% | 0 | 0 | — |
case-14 | fail→pass | 12,702 | 5,048 | -60% | 1 | 1 | 0% | 2,356 | 2,148 | -9% | 0 | 0 | — |
case-15 | fail→pass | 9,077 | 3,326 | -63% | 1 | 1 | 0% | 1,575 | 1,739 | +10% | 0 | 0 | — |
case-16 | fail→pass | 13,735 | 5,835 | -58% | 1 | 1 | 0% | 2,235 | 2,272 | +2% | 0 | 0 | — |
case-17 | pass→pass | 7,534 | 2,489 | -67% | 1 | 1 | 0% | 1,262 | 1,551 | +23% | 0 | 0 | — |
case-18 | pass→pass | 11,024 | 6,150 | -44% | 1 | 1 | 0% | 1,948 | 2,235 | +15% | 0 | 0 | — |
case-19 | fail→pass | 10,334 | 1,956 | -81% | 1 | 1 | 0% | 1,847 | 1,412 | -24% | 0 | 0 | — |
case-20 | pass→pass | 3,661 | 3,738 | +2% | 1 | 1 | 0% | 770 | 1,787 | +132% | 0 | 0 | — |
case-21 | pass→pass | 2,604 | 2,813 | +8% | 1 | 1 | 0% | 507 | 1,652 | +226% | 0 | 0 | — |
case-22 | pass→pass | 5,189 | 3,359 | -35% | 1 | 1 | 0% | 925 | 1,744 | +89% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +68 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.