Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Build in small, individually-verified increments that each leave the system working — instead of big-bang changes that fail mysteriously at the end. Use when implementing multi-part features, refactoring anything load-bearing, making large mechanical changes, or when past work produced huge diffs that were wrong somewhere unfindable. Produces the same end state as the big bang, reached through verified checkpoints you can stop at, ship from, or roll back to.
.claude/skills/mohitagw15856-incremental-implementation/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 8% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 28% | 0% |
| case-11 | ✗→✓ | ▲ Improved | -16% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 18% | 0% |
| case-14 | ✗→✓ | ▲ Improved | 8% | 0% |
The big-bang failure is always the same story: three hours of changes, then "it doesn't work", then an hour of spelunking to find WHICH of forty edits broke it. Incremental work makes the last five minutes the only suspect, always. The discipline: every increment ends with the system working and verified — not "will work once the rest lands."
| # | Increment (thin working slice) | Type | Verified by | Stoppable? | |---|---|---|---|---| | 1 | | refactor-only / behaviour | command/check] | ship / pause / rollback point |
The both-exist window (if migrating): what coexists between steps N–M, and the cutover order] Standing verification: [the command run after every increment]
(during execution, per increment: what landed → verification result → next)
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-06 | pass→fail | 13,285 | 14,540 | +9% | 1 | 1 | 0% | 2,751 | 3,995 | +45% | 0 | 0 | — |
case-01 | fail→pass | 14,676 | 9,642 | -34% | 1 | 1 | 0% | 2,807 | 3,019 | +8% | 0 | 0 | — |
case-02 | fail→fail | 15,751 | 10,812 | -31% | 1 | 1 | 0% | 2,513 | 2,930 | +17% | 0 | 0 | — |
case-03 | fail→fail | 15,882 | 12,460 | -22% | 1 | 1 | 0% | 2,821 | 3,188 | +13% | 0 | 0 | — |
case-04 | fail→fail | 3,510 | 1,637 | -53% | 1 | 1 | 0% | 692 | 1,191 | +72% | 0 | 0 | — |
case-05 | pass→pass | 19,418 | 20,447 | +5% | 1 | 1 | 0% | 4,320 | 5,544 | +28% | 0 | 0 | — |
case-07 | pass→pass | 11,703 | 12,286 | +5% | 1 | 1 | 0% | 2,137 | 3,163 | +48% | 0 | 0 | — |
case-08 | pass→pass | 14,983 | 14,028 | -6% | 1 | 1 | 0% | 2,828 | 3,340 | +18% | 0 | 0 | — |
case-09 | fail→pass | 9,930 | 6,620 | -33% | 1 | 1 | 0% | 1,619 | 2,073 | +28% | 0 | 0 | — |
case-10 | pass→pass | 13,198 | 10,568 | -20% | 1 | 1 | 0% | 2,112 | 2,505 | +19% | 0 | 0 | — |
case-11 | fail→pass | 21,719 | 12,306 | -43% | 1 | 1 | 0% | 3,490 | 2,928 | -16% | 0 | 0 | — |
case-12 | pass→pass | 17,962 | 14,598 | -19% | 1 | 1 | 0% | 2,972 | 3,406 | +15% | 0 | 0 | — |
case-13 | fail→pass | 14,038 | 11,323 | -19% | 1 | 1 | 0% | 2,269 | 2,680 | +18% | 0 | 0 | — |
case-14 | fail→pass | 14,572 | 10,659 | -27% | 1 | 1 | 0% | 2,552 | 2,746 | +8% | 0 | 0 | — |
case-15 | pass→pass | 14,352 | 11,324 | -21% | 1 | 1 | 0% | 2,591 | 3,207 | +24% | 0 | 0 | — |
case-16 | pass→pass | 13,437 | 14,297 | +6% | 1 | 1 | 0% | 2,255 | 3,084 | +37% | 0 | 0 | — |
case-17 | fail→pass | 13,733 | 8,659 | -37% | 1 | 1 | 0% | 2,568 | 2,376 | -7% | 0 | 0 | — |
case-18 | fail→pass | 15,253 | 12,680 | -17% | 1 | 1 | 0% | 2,542 | 3,113 | +22% | 0 | 0 | — |
case-19 | fail→pass | 15,131 | 13,369 | -12% | 1 | 1 | 0% | 2,734 | 3,500 | +28% | 0 | 0 | — |
case-20 | fail→fail | 18,217 | 14,887 | -18% | 1 | 1 | 0% | 3,560 | 3,697 | +4% | 0 | 0 | — |
case-21 | pass→pass | 13,252 | 7,569 | -43% | 1 | 1 | 0% | 2,150 | 2,171 | +1% | 0 | 0 | — |
case-22 | pass→pass | 13,265 | 16,962 | +28% | 1 | 1 | 0% | 2,363 | 3,903 | +65% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +32 percentage points is the difference between those two pass rates over the 22 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.