Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Output contract for substantive deliverables — outcome first, evidence-pointed reasoning, observed/derived/assumed labels, residual risk last, failures never buried, and an optional evidence-backed retrospective for explicit requests or major program closeout. Use when reporting completed work, findings, diagnoses, reviews, lane results, or a requested retrospective.
.claude/skills/happier-dev-handoff-report/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-03 | ✗→✓ | ▲ Improved | 46% | 0% |
| case-01 | ✗→✓ | ▲ Improved | 25% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 61% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 114% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 15% | 0% |
Structure any substantive report so the reader can act on the first paragraph and audit the rest. Full doctrine: docs/agent-craft.md §5 and §7.
file.ts:line, the log excerpt, the measurement, the command that produced the number. Trust should rest on checkable references, not on tone.Add a compact retrospective only when the user requests one or an approved substantial program designates it. Do not make it a routine handoff artifact.
Capture only evidence-backed learning:
One incident normally remains a local lesson. Promote it into a reusable skill or AGENTS.md only when it reflects a stable repository property, repeats across materially distinct work, or the user explicitly approves the rule. Do not collect model mix, turn counts, task counts, or other decorative metrics unless they change a future decision.
The reader acting on a misread, and caveats surfacing after the decision they should have informed.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-03 | fail→pass | 13,561 | 13,871 | +2% | 1 | 1 | 0% | 2,235 | 3,264 | +46% | 0 | 0 | — |
case-01 | fail→pass | 15,240 | 13,732 | -10% | 1 | 1 | 0% | 2,428 | 3,034 | +25% | 0 | 0 | — |
case-02 | fail→fail | 13,472 | 15,110 | +12% | 1 | 1 | 0% | 2,117 | 3,137 | +48% | 0 | 0 | — |
case-04 | pass→fail | 7,863 | 13,771 | +75% | 1 | 1 | 0% | 1,371 | 3,009 | +119% | 0 | 0 | — |
case-05 | pass→pass | 15,970 | 16,977 | +6% | 1 | 1 | 0% | 2,780 | 3,546 | +28% | 0 | 0 | — |
case-06 | pass→pass | 11,794 | 12,721 | +8% | 1 | 1 | 0% | 2,148 | 3,017 | +40% | 0 | 0 | — |
case-07 | fail→pass | 6,852 | 8,267 | +21% | 1 | 1 | 0% | 1,375 | 2,217 | +61% | 0 | 0 | — |
case-08 | fail→pass | 7,110 | 11,311 | +59% | 1 | 1 | 0% | 1,275 | 2,734 | +114% | 0 | 0 | — |
case-09 | pass→pass | 7,871 | 8,131 | +3% | 1 | 1 | 0% | 1,556 | 2,232 | +43% | 0 | 0 | — |
case-10 | pass→pass | 13,298 | 7,091 | -47% | 1 | 1 | 0% | 2,243 | 2,041 | -9% | 0 | 0 | — |
case-11 | fail→pass | 12,095 | 9,409 | -22% | 1 | 1 | 0% | 2,107 | 2,418 | +15% | 0 | 0 | — |
case-12 | pass→pass | 11,339 | 13,027 | +15% | 1 | 1 | 0% | 1,913 | 3,044 | +59% | 0 | 0 | — |
case-13 | pass→pass | 5,855 | 7,301 | +25% | 1 | 1 | 0% | 956 | 2,135 | +123% | 0 | 0 | — |
case-14 | fail→pass | 11,358 | 11,108 | -2% | 1 | 1 | 0% | 1,821 | 2,522 | +38% | 0 | 0 | — |
case-15 | fail→fail | 9,934 | 14,062 | +42% | 1 | 1 | 0% | 1,664 | 3,166 | +90% | 0 | 0 | — |
case-16 | pass→pass | 14,389 | 7,979 | -45% | 1 | 1 | 0% | 2,321 | 2,236 | -4% | 0 | 0 | — |
case-17 | fail→fail | 7,428 | 10,656 | +43% | 1 | 1 | 0% | 1,244 | 2,334 | +88% | 0 | 0 | — |
case-18 | fail→pass | 7,963 | 8,095 | +2% | 1 | 1 | 0% | 1,259 | 2,111 | +68% | 0 | 0 | — |
case-19 | fail→pass | 9,973 | 10,766 | +8% | 1 | 1 | 0% | 1,610 | 2,624 | +63% | 0 | 0 | — |
case-20 | pass→pass | 8,651 | 8,048 | -7% | 1 | 1 | 0% | 1,587 | 2,236 | +41% | 0 | 0 | — |
case-21 | pass→pass | 7,030 | 6,116 | -13% | 1 | 1 | 0% | 1,208 | 1,728 | +43% | 0 | 0 | — |
case-22 | pass→pass | 7,723 | 8,038 | +4% | 1 | 1 | 0% | 1,311 | 2,152 | +64% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +32 percentage points is the difference between those two pass rates over the 22 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.