Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Audits an existing Clean Room run and steers it back to missed gates without expanding declared scope.
.claude/skills/hashgraph-online-refocus/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-10 | ✗→✓ | ▲ Improved | 16% | 0% |
| case-04 | ✗→✓ | ▲ Improved | -2% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 99% | 0% |
| case-06 | ✗→✓ | ▲ Improved | -20% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 75% | 0% |
Refocus realigns the current run to the declared scope, controller policy, artifact schemas, and clean-room boundary.
Refocus does not optimize, expand, or reinterpret the task. It does not invent new requirements or add behavior beyond preflight-goal.json, task-manifest.json, clean-run-context.json, ledgers, implementation plan/report, QC, and abstract delta tickets.
Use the canonical clean-room skill workflow and references in this plugin. Preserve the same clean-room boundary, role separation, artifact schemas, leakage rules, implementation-root rules, and hook expectations.
Compare current artifacts to the canonical gate checklist:
preflight-goal.schema.json, and is referenced by task-manifest.json with preflight_goal_ref and preflight_goal_sha256.handoff_sequence and does not skip Stage 0.init-config.json drift is reported instead of silently applied.clean-run-context.json exists before clean roles run and excludes source roots, visual roots, contaminated roots, source index refs, visual index refs, and ledger paths.clean-run-context.json records artifact-only coordination: Agent 0 does not directly steer Agent 2, Agent 3, or Agent 4, and clean implementation/polish roles report to Agent 0 only at terminal status.role-session-brief.json inside the recorded budgets. controller-status.json remains contaminated-side only.task-manifest.json units.task-manifest.json, source-index.json, visual-index.json, raw screenshots, source or visual paths, raw diffs, copied comments, copied visible words, private identifiers, exact UI palettes/layouts/iconography, source-shaped pseudocode, and contaminated ledgers.implementation-plan.json when the run reached that gate.implementation-report.json when the run reached implementation.qc-report.json with schema, leakage, coverage, and abstract delta ticket status when the run reached that gate.polish-report.json when the run reached final polish review.Validate schemas and handoff hashes before trusting the artifacts. Use source-index.json or visual-index.json only on the contaminated side and only when referenced by the task manifest.
Emit missed-gate findings only:
not verified unless clean-room-skill run --dry-run (or npx clean-room-skill@latest run --dry-run if the binary is not available) succeeds against the canonical task-manifest.json.public_contract_refs, terminal implementation reports, and coverage-ledger public_surface_coverage.Do not suggest speculative improvements. Do not change source scope, target profile, public API, or implementation plan. If the user asks to add scope, stop and route to a new scope gate instead of silently expanding the run.
Return a bounded next-action plan containing:
If root separation, authorization, or clean/contaminated boundary cannot be proven, the single next action is to stop and repair that gate.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-10 | fail→pass | 14,123 | 12,870 | -9% | 1 | 1 | 0% | 2,142 | 2,484 | +16% | 0 | 0 | — |
case-04 | fail→pass | 20,959 | 13,766 | -34% | 1 | 1 | 0% | 2,623 | 2,559 | -2% | 0 | 0 | — |
case-01 | fail→fail | 13,267 | 17,446 | +31% | 1 | 1 | 0% | 1,058 | 1,474 | +39% | 0 | 0 | — |
case-02 | fail→pass | 12,155 | 32,473 | +167% | 1 | 1 | 0% | 1,415 | 2,813 | +99% | 0 | 0 | — |
case-03 | fail→fail | 10,732 | 21,259 | +98% | 1 | 1 | 0% | 1,692 | 1,375 | -19% | 0 | 0 | — |
case-05 | fail→fail | 13,388 | 13,907 | +4% | 1 | 1 | 0% | 2,076 | 2,696 | +30% | 0 | 0 | — |
case-06 | fail→pass | 18,928 | 12,329 | -35% | 1 | 1 | 0% | 2,814 | 2,239 | -20% | 0 | 0 | — |
case-07 | fail→pass | 14,429 | 6,590 | -54% | 1 | 1 | 0% | 1,375 | 2,402 | +75% | 0 | 0 | — |
case-08 | fail→pass | 25,575 | 14,934 | -42% | 1 | 1 | 0% | 628 | 2,846 | +353% | 0 | 0 | — |
case-09 | fail→pass | 20,027 | 12,053 | -40% | 1 | 1 | 0% | 2,338 | 2,241 | -4% | 0 | 0 | — |
case-11 | pass→pass | 15,236 | 6,534 | -57% | 1 | 1 | 0% | 2,262 | 2,159 | -5% | 0 | 0 | — |
case-12 | fail→pass | 37,863 | 15,383 | -59% | 1 | 1 | 0% | 1,021 | 2,992 | +193% | 0 | 0 | — |
case-13 | fail→pass | 13,509 | 7,856 | -42% | 1 | 1 | 0% | 1,442 | 2,413 | +67% | 0 | 0 | — |
case-14 | fail→pass | 21,006 | 7,429 | -65% | 1 | 1 | 0% | 2,265 | 2,569 | +13% | 0 | 0 | — |
case-15 | fail→pass | 15,184 | 9,958 | -34% | 1 | 1 | 0% | 2,516 | 2,259 | -10% | 0 | 0 | — |
case-16 | fail→pass | 10,905 | 24,919 | +129% | 1 | 1 | 0% | 887 | 2,810 | +217% | 0 | 0 | — |
case-17 | fail→pass | 12,690 | 7,835 | -38% | 1 | 1 | 0% | 1,920 | 2,450 | +28% | 0 | 0 | — |
case-18 | fail→pass | 16,399 | 12,601 | -23% | 1 | 1 | 0% | 2,477 | 2,479 | +0% | 0 | 0 | — |
case-19 | fail→pass | 20,813 | 13,743 | -34% | 1 | 1 | 0% | 2,458 | 2,759 | +12% | 0 | 0 | — |
case-20 | pass→pass | 15,515 | 26,355 | +70% | 1 | 1 | 0% | 1,724 | 5,370 | +211% | 0 | 0 | — |
case-21 | pass→pass | 10,664 | 17,322 | +62% | 1 | 1 | 0% | 2,038 | 3,334 | +64% | 0 | 0 | — |
case-22 | pass→fail | 14,837 | 12,010 | -19% | 1 | 1 | 0% | 1,587 | 3,207 | +102% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 18 counted toward the lift figure. The other 4 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +64 percentage points is the difference between those two pass rates over the 18 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.