Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Create or refine PRODUCT.md while separating
.claude/skills/boshu2-product/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | -21% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 207% | 0% |
| case-06 | ✗→✓ | ▲ Improved | -23% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 31% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 39% | 0% |
Create or refine a product contract with the user's authority.
The contract works because it forces every claim into evidence or labeled aspiration; if a reader cannot tell which is which, the document has failed.
this repository's own product surfaces, start from the canonical category in docs/contracts/ubiquitous-language.md (AgentOps is the operations layer for agentic engineering) and preserve the small ownership boundary — never re-frame AgentOps as an execution orchestrator, factory, or loop.
claim under ## Proven, each carrying a resolvable citation (a README, release, test, or evidence source). Put unproven hopes, evidence gaps, and not-yet-measured success signals under ## Assumptions. A claim that cannot cite a source belongs under ## Assumptions, never ## Proven.
Named failure mode — aspiration laundering: an unproven hope written in the ## Proven section; once laundered, every downstream plan inherits a false premise. Detector: any ## Proven claim without a resolvable citation is a laundered aspiration — move it to ## Assumptions or cite it.
Anti-pattern: rewriting a healthy PRODUCT.md wholesale because the session has fresh opinions. Corrective: refine only the sections the user asked to change and preserve the rest byte-for-byte.
Product does not select work, invoke a loop, repair itself, validate itself, or choose delivery.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 19,146 | 4,948 | -74% | 1 | 1 | 0% | 3,075 | 751 | -76% | 0 | 0 | — |
case-02 | fail→fail | 3,677 | 5,849 | +59% | 1 | 1 | 0% | 582 | 678 | +16% | 0 | 0 | — |
case-03 | fail→fail | 16,264 | 6,083 | -63% | 1 | 1 | 0% | 2,439 | 833 | -66% | 0 | 0 | — |
case-04 | fail→pass | 21,618 | 13,347 | -38% | 1 | 1 | 0% | 3,634 | 2,874 | -21% | 0 | 0 | — |
case-05 | fail→pass | 3,899 | 8,704 | +123% | 1 | 1 | 0% | 634 | 1,949 | +207% | 0 | 0 | — |
case-06 | fail→pass | 12,744 | 5,888 | -54% | 1 | 1 | 0% | 1,792 | 1,383 | -23% | 0 | 0 | — |
case-07 | fail→pass | 10,017 | 10,163 | +1% | 1 | 1 | 0% | 1,580 | 2,062 | +31% | 0 | 0 | — |
case-08 | fail→fail | 2,510 | 5,837 | +133% | 1 | 1 | 0% | 371 | 1,321 | +256% | 0 | 0 | — |
case-09 | fail→pass | 8,339 | 7,756 | -7% | 1 | 1 | 0% | 1,163 | 1,617 | +39% | 0 | 0 | — |
case-10 | fail→pass | 22,504 | 14,201 | -37% | 1 | 1 | 0% | 3,943 | 2,471 | -37% | 0 | 0 | — |
case-11 | pass→pass | 8,902 | 4,669 | -48% | 1 | 1 | 0% | 1,503 | 1,181 | -21% | 0 | 0 | — |
case-12 | fail→pass | 7,364 | 3,322 | -55% | 1 | 1 | 0% | 1,129 | 989 | -12% | 0 | 0 | — |
case-13 | pass→pass | 9,661 | 7,684 | -20% | 1 | 1 | 0% | 1,611 | 1,679 | +4% | 0 | 0 | — |
case-14 | fail→pass | 12,451 | 10,839 | -13% | 1 | 1 | 0% | 1,986 | 2,231 | +12% | 0 | 0 | — |
case-15 | fail→pass | 12,018 | 8,216 | -32% | 1 | 1 | 0% | 1,823 | 1,664 | -9% | 0 | 0 | — |
case-16 | fail→fail | 10,187 | 4,661 | -54% | 1 | 1 | 0% | 1,619 | 1,223 | -24% | 0 | 0 | — |
case-17 | pass→pass | 17,097 | 11,647 | -32% | 1 | 1 | 0% | 2,652 | 2,254 | -15% | 0 | 0 | — |
case-18 | fail→pass | 15,941 | 14,580 | -9% | 1 | 1 | 0% | 2,493 | 2,715 | +9% | 0 | 0 | — |
case-19 | fail→pass | 8,841 | 5,764 | -35% | 1 | 1 | 0% | 1,383 | 1,440 | +4% | 0 | 0 | — |
case-20 | pass→pass | 3,187 | 6,417 | +101% | 1 | 1 | 0% | 526 | 1,426 | +171% | 0 | 0 | — |
case-21 | pass→pass | 3,244 | 7,095 | +119% | 1 | 1 | 0% | 425 | 1,472 | +246% | 0 | 0 | — |
case-22 | pass→fail | 7,202 | 7,309 | +1% | 1 | 1 | 0% | 1,165 | 960 | -18% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 18 counted toward the lift figure. The other 4 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +45 percentage points is the difference between those two pass rates over the 18 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.