Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Verification entrypoint for repo-harness workflow readiness. Runs workflow gates, task sync, contract checks, inspector, and migration dry-run before merge or release.
.claude/skills/ancienttwo-repo-harness-check/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-17 | ✗→✓ | ▲ Improved | -8% | 0% |
| case-06 | ✗→✓ | ▲ Improved | -18% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 6% | 0% |
| case-08 | ✗→✓ | ▲ Improved | -12% | 0% |
| case-09 | ✗→✓ | ▲ Improved | -9% | 0% |
Use this command when the user asks whether the harness, migration, or release surface is ready.
CLAUDE.md for Claude, AGENTS.md for Codex) ## Required Checks section, then apply its scope conditions and run the applicable commands through the global/package helper runtime. This self-host source repo may also use root scripts/ for source-only maintenance commands. Use focused checks by default; a conditional full-suite command is not an unconditional requirement. Consume current canonical acceptance evidence for already-satisfied expensive criteria. After a bounded follow-up edit, have the parent retain the baseline full-run identity and revise final criteria to focused delta checks when no explicit full-suite requirement or uncovered integration risk remains; never relabel a baseline pass as current-subject full evidence. If ## Required Checks is missing or empty, report that as the first blocking finding instead of substituting a default list.repo-harness run check-agent-tooling --host both --jsonhealth/check/mermaid as hard failures.effectiveness:
bun run benchmark:skills --eval <slug> withfull_test_count > 0, dry_run_ratio <= 30%, and graders reported
A file-coupled contract-run worker prepares executable evidence through verify-sprint --prepare-acceptance. Inspect the immutable run artifact and its current subject, contract, target revision, toolchain context, and per-criterion results. Valid executed and exact-context reused passes count. Return missing, stale, or failed criteria to the canonical runner; do not independently launch the same suite or accept transcript assertions as evidence.
For a bugfix contract, confirm ## Root Cause Evidence states a testable root_cause, a working repro, a regression_guard that also appears under exit_criteria.tests_pass, and a pre_fix_failure_artifact showing a non-zero PRE_FIX_EXIT= line for that guard, not a passing run. Also check whether task_profile was mislabeled or left out entirely: an omitted task_profile defaults to legacy pass-through (non-bugfix) by design, so confirm that default is actually correct here rather than an evasion of the gate.
## Required Checks section is the single source of truth.| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-17 | fail→pass | 12,917 | 6,176 | -52% | 1 | 1 | 0% | 2,015 | 1,850 | -8% | 0 | 0 | — |
case-01 | fail→fail | 10,406 | 4,184 | -60% | 1 | 1 | 0% | 308 | 1,192 | +287% | 0 | 0 | — |
case-02 | fail→fail | 3,531 | 4,425 | +25% | 1 | 1 | 0% | 482 | 1,178 | +144% | 0 | 0 | — |
case-03 | fail→fail | 4,626 | 8,637 | +87% | 1 | 1 | 0% | 331 | 1,109 | +235% | 0 | 0 | — |
case-04 | pass→pass | 10,982 | 4,135 | -62% | 1 | 1 | 0% | 1,711 | 1,509 | -12% | 0 | 0 | — |
case-05 | pass→pass | 7,531 | 3,347 | -56% | 1 | 1 | 0% | 1,155 | 1,355 | +17% | 0 | 0 | — |
case-06 | fail→pass | 10,428 | 3,134 | -70% | 1 | 1 | 0% | 1,500 | 1,226 | -18% | 0 | 0 | — |
case-07 | fail→pass | 8,185 | 3,467 | -58% | 1 | 1 | 0% | 1,167 | 1,242 | +6% | 0 | 0 | — |
case-08 | fail→pass | 11,798 | 3,919 | -67% | 1 | 1 | 0% | 1,588 | 1,401 | -12% | 0 | 0 | — |
case-09 | fail→pass | 11,049 | 3,801 | -66% | 1 | 1 | 0% | 1,368 | 1,251 | -9% | 0 | 0 | — |
case-10 | fail→pass | 12,191 | 4,428 | -64% | 1 | 1 | 0% | 1,849 | 1,550 | -16% | 0 | 0 | — |
case-11 | pass→pass | 11,193 | 3,940 | -65% | 1 | 1 | 0% | 1,689 | 1,282 | -24% | 0 | 0 | — |
case-12 | pass→pass | 12,393 | 3,863 | -69% | 1 | 1 | 0% | 1,910 | 1,332 | -30% | 0 | 0 | — |
case-13 | fail→pass | 9,757 | 4,622 | -53% | 1 | 1 | 0% | 1,459 | 1,080 | -26% | 0 | 0 | — |
case-14 | pass→pass | 8,110 | 4,417 | -46% | 1 | 1 | 0% | 1,285 | 1,416 | +10% | 0 | 0 | — |
case-15 | fail→pass | 16,047 | 3,368 | -79% | 1 | 1 | 0% | 2,527 | 1,271 | -50% | 0 | 0 | — |
case-16 | fail→pass | 9,492 | 3,810 | -60% | 1 | 1 | 0% | 1,197 | 1,398 | +17% | 0 | 0 | — |
case-18 | fail→pass | 12,131 | 4,010 | -67% | 1 | 1 | 0% | 1,835 | 1,495 | -19% | 0 | 0 | — |
case-19 | pass→pass | 9,112 | 3,991 | -56% | 1 | 1 | 0% | 1,285 | 1,394 | +8% | 0 | 0 | — |
case-20 | pass→pass | 12,131 | 7,977 | -34% | 1 | 1 | 0% | 1,716 | 2,047 | +19% | 0 | 0 | — |
case-21 | fail→pass | 12,895 | 4,492 | -65% | 1 | 1 | 0% | 1,901 | 1,514 | -20% | 0 | 0 | — |
case-22 | pass→fail | 10,352 | 9,067 | -12% | 1 | 1 | 0% | 1,403 | 2,172 | +55% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 20 counted toward the lift figure. The other 2 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +45 percentage points is the difference between those two pass rates over the 20 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
The publisher has shipped newer versions since this run, so these numbers describe v1, not the version currently listed.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.