Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Implements a coding task from a written spec with Done-When acceptance items. Use when given a spec to implement: stub the interface, implement it, then write tests that validate the Done-When items and any important behavior beyond them.
.claude/skills/ryanthedev-build-doctrine/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-03 | ✗→✓ | ▲ Improved | -43% | 0% |
| case-04 | ✗→✓ | ▲ Improved | -20% | 0% |
| case-10 | ✗→✓ | ▲ Improved | -36% | 0% |
| case-15 | ✗→✓ | ▲ Improved | -49% | 0% |
| case-17 | ✗→✓ | ▲ Improved | -41% | 0% |
Implement the task described in the provided spec so it satisfies the spec's Done-When (DW) items. Write all code and tests to the exact output paths the spec names.
Test-driven, working red-green, one behavior at a time:
Cover every DW item — that is the floor, not the ceiling. Then go past the DW list: write failing tests first for the edge cases, error paths, and boundaries you anticipate too, even when no DW item names them. Name DW-item tests for their DW-ID (e.g. test_DW_1_1_parses_hours).
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | pass→pass | 5,840 | 2,970 | -49% | 1 | 1 | 0% | 968 | 635 | -34% | 0 | 0 | — |
case-02 | pass→pass | 8,565 | 4,071 | -52% | 1 | 1 | 0% | 1,291 | 897 | -31% | 0 | 0 | — |
case-03 | fail→pass | 9,506 | 3,400 | -64% | 1 | 1 | 0% | 1,342 | 767 | -43% | 0 | 0 | — |
case-04 | fail→pass | 3,933 | 2,059 | -48% | 1 | 1 | 0% | 656 | 524 | -20% | 0 | 0 | — |
case-05 | pass→pass | 10,302 | 3,852 | -63% | 1 | 1 | 0% | 1,583 | 677 | -57% | 0 | 0 | — |
case-06 | pass→pass | 10,027 | 3,864 | -61% | 1 | 1 | 0% | 1,587 | 802 | -49% | 0 | 0 | — |
case-07 | pass→pass | 4,349 | 3,041 | -30% | 1 | 1 | 0% | 732 | 631 | -14% | 0 | 0 | — |
case-08 | pass→pass | 7,901 | 2,984 | -62% | 1 | 1 | 0% | 1,287 | 614 | -52% | 0 | 0 | — |
case-09 | pass→pass | 12,095 | 6,015 | -50% | 1 | 1 | 0% | 1,773 | 1,093 | -38% | 0 | 0 | — |
case-10 | fail→pass | 4,415 | 2,234 | -49% | 1 | 1 | 0% | 786 | 503 | -36% | 0 | 0 | — |
case-11 | pass→pass | 9,490 | 3,331 | -65% | 1 | 1 | 0% | 1,294 | 636 | -51% | 0 | 0 | — |
case-12 | pass→pass | 9,037 | 3,808 | -58% | 1 | 1 | 0% | 1,369 | 733 | -46% | 0 | 0 | — |
case-13 | pass→pass | 6,697 | 3,554 | -47% | 1 | 1 | 0% | 1,050 | 722 | -31% | 0 | 0 | — |
case-14 | pass→pass | 2,036 | 1,832 | -10% | 1 | 1 | 0% | 267 | 474 | +78% | 0 | 0 | — |
case-15 | fail→pass | 16,668 | 4,079 | -76% | 1 | 1 | 0% | 1,712 | 868 | -49% | 0 | 0 | — |
case-16 | pass→pass | 10,482 | 2,411 | -77% | 1 | 1 | 0% | 1,679 | 615 | -63% | 0 | 0 | — |
case-17 | fail→pass | 5,792 | 2,448 | -58% | 1 | 1 | 0% | 904 | 531 | -41% | 0 | 0 | — |
case-18 | pass→pass | 8,737 | 3,608 | -59% | 1 | 1 | 0% | 1,416 | 750 | -47% | 0 | 0 | — |
case-19 | pass→pass | 15,446 | 2,964 | -81% | 1 | 1 | 0% | 1,719 | 642 | -63% | 0 | 0 | — |
case-20 | pass→pass | 4,416 | 3,190 | -28% | 1 | 1 | 0% | 729 | 651 | -11% | 0 | 0 | — |
case-21 | pass→pass | 2,652 | 4,511 | +70% | 1 | 1 | 0% | 321 | 912 | +184% | 0 | 0 | — |
case-22 | pass→pass | 6,270 | 4,065 | -35% | 1 | 1 | 0% | 901 | 691 | -23% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +23 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.