Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Route explicit repo-harness setup, planning, execution, verification, and handoff actions through deterministic repository state.
.claude/skills/ancienttwo-repo-harness/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 7% | 0% |
| case-03 | ✗→✓ | ▲ Improved | -46% | 0% |
| case-08 | ✗→✓ | ▲ Improved | -19% | 0% |
| case-09 | ✗→✓ | ▲ Improved | -9% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 12% | 0% |
This is an evaluation-only delta over the fixture's minimum-effective-interview baseline. It has no authority to approve or implement a plan.
Use the treatment only when explicitly requested or when a case changes high-risk architecture, data authority or ownership, security, permissions, money, deletion, concurrency, recovery semantics, or a hard-to-reverse public interface. Bypass it for small bug fixes, decision-complete acceptance criteria, documentation, formatting, and renames.
CASE.md; do not ask the user for known facts.the current frontier only after every prerequisite is resolved.
recommended default and the effect of every option.
continue; unvisited branches remain explicit.
[UNKNOWN:BLOCKING] and keep theplan Draft.
docs/spec.md#Canonical Terms; user need/non-goal to PRD; boundary and trade-off to Plan; allowed scope and failure semantics to Plan + Contract; verifiable behavior to Contract exit_criteria; long-lived architecture to its architecture module/request. Reversible defaults use [ASSUMED].
plan Approved and never start implementation from this mode.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 5,581 | 38,306 | +586% | 1 | 1 | 0% | 267 | 6,538 | +2349% | 0 | 0 | — |
case-02 | fail→pass | 33,963 | 31,551 | -7% | 1 | 1 | 0% | 5,368 | 5,721 | +7% | 0 | 0 | — |
case-03 | fail→pass | 29,277 | 18,051 | -38% | 1 | 1 | 0% | 5,366 | 2,874 | -46% | 0 | 0 | — |
case-04 | pass→pass | 11,804 | 65,337 | +454% | 1 | 1 | 0% | 2,093 | 1,484 | -29% | 0 | 0 | — |
case-05 | pass→pass | 10,580 | 15,116 | +43% | 1 | 1 | 0% | 1,810 | 1,893 | +5% | 0 | 0 | — |
case-06 | fail→fail | 16,201 | 4,473 | -72% | 1 | 1 | 0% | 1,354 | 1,038 | -23% | 0 | 0 | — |
case-07 | pass→pass | 28,184 | 15,915 | -44% | 1 | 1 | 0% | 4,261 | 2,805 | -34% | 0 | 0 | — |
case-08 | fail→pass | 20,667 | 14,962 | -28% | 1 | 1 | 0% | 2,893 | 2,344 | -19% | 0 | 0 | — |
case-09 | fail→pass | 65,938 | 20,249 | -69% | 1 | 1 | 0% | 3,718 | 3,388 | -9% | 0 | 0 | — |
case-10 | fail→pass | 17,919 | 20,147 | +12% | 1 | 1 | 0% | 2,793 | 3,138 | +12% | 0 | 0 | — |
case-11 | fail→pass | 26,692 | 22,946 | -14% | 1 | 1 | 0% | 4,216 | 3,741 | -11% | 0 | 0 | — |
case-12 | fail→pass | 25,904 | 27,661 | +7% | 1 | 1 | 0% | 3,327 | 3,970 | +19% | 0 | 0 | — |
case-13 | pass→pass | 34,883 | 31,501 | -10% | 1 | 1 | 0% | 4,126 | 5,434 | +32% | 0 | 0 | — |
case-14 | fail→pass | 142,326 | 18,822 | -87% | 1 | 1 | 0% | 9,354 | 3,132 | -67% | 0 | 0 | — |
case-15 | fail→pass | 8,043 | 39,371 | +390% | 1 | 1 | 0% | 164 | 6,319 | +3753% | 0 | 0 | — |
case-16 | fail→pass | 37,153 | 21,908 | -41% | 1 | 1 | 0% | 5,651 | 4,076 | -28% | 0 | 0 | — |
case-17 | fail→pass | 21,336 | 26,473 | +24% | 1 | 1 | 0% | 3,300 | 3,518 | +7% | 0 | 0 | — |
case-18 | pass→pass | 36,091 | 25,834 | -28% | 1 | 1 | 0% | 5,617 | 4,724 | -16% | 0 | 0 | — |
case-19 | fail→pass | 26,579 | 13,070 | -51% | 1 | 1 | 0% | 3,944 | 1,364 | -65% | 0 | 0 | — |
case-20 | fail→pass | 30,050 | 28,459 | -5% | 1 | 1 | 0% | 5,555 | 2,712 | -51% | 0 | 0 | — |
case-21 | fail→fail | 21,792 | 29,402 | +35% | 1 | 1 | 0% | 3,402 | 3,623 | +6% | 0 | 0 | — |
case-22 | fail→pass | 63,968 | 4,738 | -93% | 1 | 1 | 0% | 6,711 | 896 | -87% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 20 counted toward the lift figure. The other 2 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +64 percentage points is the difference between those two pass rates over the 20 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.