Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Execute one bounded RED to GREEN experiment
.claude/skills/boshu2-implement/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-05 | ✗→✓ | ▲ Improved | -8% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 179% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 41% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 106% | 0% |
| case-14 | ✗→✓ | ▲ Improved | 88% | 0% |
Execute exactly one bounded experiment described by the resolved bead or caller intent. Implement owns subject edits and factual evidence. It does not create a second planning record or a model-authored candidate packet.
may snapshot and hash that source automatically for drift detection.
applies only when acceptance is behavioral: preserve evidence that the check fails for the expected missing behavior. Relocations, doc merges, and pure refactors need no failing-check ritual — record an honest green pre-change baseline instead.
acceptance test.
subject-manifest.v1 fromthe before/after subject. Do not make the model transcribe those facts.
response or runtime channel. Stop.
Specialists such as standards, domain, test, refactor, and security may provide advice. They are never hard dependencies and cannot add lifecycle authority.
During edits, run the smallest deterministic checks that can falsify the active change. Reuse exact-input receipts when their subject and tool identity still match. Run an expensive full-suite check at the integration boundary, or earlier only when the intent explicitly makes it the first acceptance check. Repeatedly replaying the full suite after every focused edit adds latency, not proof.
On discovering a live consumer of the change outside the declared write scope — a test asserting the old path, a generated twin, a gate reading the moved file — stop and report the exact file and line to the caller. Do not silently expand scope to absorb it or revise the intent from Implement. The caller may revise the source intent and start a separate invocation.
Before declaring GREEN, self-audit the diff for mocks, placeholders, TODO stubs, hardcoded fixture values, weakened assertions, regenerated goldens, widened tolerances, suppression directives, or specification edits standing in for real behavior. When the diff changes a test, gate, fixture, golden, or acceptance source, state why the original intent requires that change and confirm that green came from the implemented behavior rather than a weakened oracle. A check that passes against a substitute or weakened oracle is not evidence for the acceptance criterion; either finish the behavior or report it as not built.
semantic validator.
intent for a caller to start separately.
or validation loop.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-05 | fail→pass | 8,998 | 2,898 | -68% | 1 | 1 | 0% | 1,302 | 1,193 | -8% | 0 | 0 | — |
case-01 | fail→fail | 8,047 | 4,216 | -48% | 1 | 1 | 0% | 313 | 828 | +165% | 0 | 0 | — |
case-02 | fail→fail | 3,891 | 5,423 | +39% | 1 | 1 | 0% | 226 | 874 | +287% | 0 | 0 | — |
case-03 | fail→fail | 5,027 | 4,382 | -13% | 1 | 1 | 0% | 235 | 829 | +253% | 0 | 0 | — |
case-04 | fail→fail | 5,151 | 5,800 | +13% | 1 | 1 | 0% | 284 | 901 | +217% | 0 | 0 | — |
case-06 | fail→pass | 3,511 | 4,116 | +17% | 1 | 1 | 0% | 499 | 1,391 | +179% | 0 | 0 | — |
case-07 | pass→pass | 7,731 | 4,457 | -42% | 1 | 1 | 0% | 1,301 | 1,474 | +13% | 0 | 0 | — |
case-08 | fail→pass | 6,492 | 5,142 | -21% | 1 | 1 | 0% | 1,075 | 1,511 | +41% | 0 | 0 | — |
case-09 | pass→pass | 4,181 | 2,662 | -36% | 1 | 1 | 0% | 650 | 1,090 | +68% | 0 | 0 | — |
case-10 | pass→pass | 4,951 | 3,608 | -27% | 1 | 1 | 0% | 717 | 1,219 | +70% | 0 | 0 | — |
case-11 | pass→pass | 7,655 | 3,906 | -49% | 1 | 1 | 0% | 1,154 | 1,157 | +0% | 0 | 0 | — |
case-12 | fail→pass | 4,078 | 3,754 | -8% | 1 | 1 | 0% | 639 | 1,314 | +106% | 0 | 0 | — |
case-13 | pass→fail | 6,675 | 6,735 | +1% | 1 | 1 | 0% | 1,066 | 1,031 | -3% | 0 | 0 | — |
case-14 | fail→pass | 5,751 | 5,711 | -1% | 1 | 1 | 0% | 845 | 1,587 | +88% | 0 | 0 | — |
case-15 | pass→pass | 5,555 | 6,915 | +24% | 1 | 1 | 0% | 927 | 1,908 | +106% | 0 | 0 | — |
case-16 | pass→pass | 10,685 | 4,887 | -54% | 1 | 1 | 0% | 1,595 | 1,397 | -12% | 0 | 0 | — |
case-17 | pass→fail | 5,337 | 2,756 | -48% | 1 | 1 | 0% | 858 | 1,047 | +22% | 0 | 0 | — |
case-18 | fail→pass | 10,419 | 3,088 | -70% | 1 | 1 | 0% | 1,637 | 1,119 | -32% | 0 | 0 | — |
case-19 | fail→pass | 10,524 | 5,468 | -48% | 1 | 1 | 0% | 1,801 | 1,624 | -10% | 0 | 0 | — |
case-20 | pass→fail | 25,165 | 3,517 | -86% | 1 | 1 | 0% | 4,178 | 866 | -79% | 0 | 0 | — |
case-21 | fail→fail | 8,282 | 4,799 | -42% | 1 | 1 | 0% | 1,321 | 1,343 | +2% | 0 | 0 | — |
case-22 | pass→fail | 14,232 | 4,828 | -66% | 1 | 1 | 0% | 2,264 | 909 | -60% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 15 counted toward the lift figure. The other 7 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +14 percentage points is the difference between those two pass rates over the 15 comparable cases. 4 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.