Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Run one explicitly authorized, evidence-backed improvement to a repository's agent guidance, tools, runbooks, or validation. Use only when the user invokes `$improve-harness` or explicitly asks to improve the Harness after observed reusable agent friction. Do not use for ordinary product changes, speculative cleanup, one unexplained agent mistake, or automatic post-task reflection.
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-07 | ✗→✓ | ▲ Improved | 102% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 9% | 0% |
| case-15 | ✗→✓ | ▲ Improved | -34% | 0% |
| case-16 | ✗→✓ | ▲ Improved | -18% | 0% |
| case-17 | ✗→✓ | ▲ Improved | 14% | 0% |
Improve one bounded future-agent behavior without turning every difficult task into permanent process. Keep consumer truth with its owner and require a fresh rerun before claiming improvement.
AGENTS.md, docs/WORKFLOW.md, and applicable local instructions.request to report friction does not authorize edits.
changes. Preserve all existing work.
product policy, weakening proof, adding credentials, or mutating external systems.
Use an observed task trajectory when available. Record:
authority; and
Do not diagnose a worker limitation from one run. If no observed baseline exists, stop with an experiment proposal; do not manufacture one.
Copy docs/templates/harness-improvement.md to docs/plans/active/harness-improvement-<slug>.md. Reuse an existing active record for the same experiment.
Trace the failure upstream to the first owner that could have prevented or exposed it:
verification failed.
the invariant.
Assign the correction to repository-harness, the consumer repository, the external environment, or a human decision. Do not copy consumer commands or policy into a generic upstream template.
Before editing, write:
textIf <smallest change> is added at <owner>, then a fresh agent will <observable change> on <representative job>, because <mechanism>. Evidence that would weaken this: Maintenance owner and removal condition:
Make only the authorized intervention. Prefer an existing owner, a clearer route, an actionable diagnostic, a runbook fact, a type or API, or claim-matched proof over a parallel framework. Keep unknown policy unknown. Run repository-native checks that protect the changed boundary.
Use a fresh agent session and an equivalent starting state. Hold the worker, task class, authority, tools, and relevant external conditions materially steady. Record separately whether the intervention was available, retrieved or invoked, and relevant.
If a fresh rerun is not authorized or available, leave the record active with Decision: pending fresh rerun. Report the exact next task; do not claim the Harness improved.
Compare accepted outcome, claim-matched proof, human intervention, retries, authority behavior, and maintenance cost.
job enough to justify its cost.
to use.
the job.
Record the decision, evidence, owner, and removal condition. Move the record to docs/plans/completed/ only after native validation and the fresh-rerun decision. Preserve a removed intervention's result in the completed record.
Return:
keep, revise, remove, or pending fresh rerun; andOther measured skills in the registry, with their headline benchmark lift.