Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Ramps implementation ambition a notch only after the prior increment is understood. Use when building a feature you must understand, not just ship.
.claude/skills/athola-graduated-implementation/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 34% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 23% | 0% |
| case-15 | ✗→✓ | ▲ Improved | 34% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 24% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 37% | 0% |
> Start with the smallest slice you can fully understand. Earn the > next notch by proving you understood the last one. Ambition that > outruns understanding is how a fluent diff becomes an unverifiable > one.
The sibling skill imbue:assisted-mastery fades scaffolding as competence grows. This skill ramps the other axis: the ambition of the next increment. They are the two directions of one move, the graduated practice that turned novices into experts long before agents existed. Not "ban the tool," but "couple the next challenge to demonstrated competence on the last one."
The learning sciences give the move a number. Wilson et al. (2019, Nature Communications 10:4646) derive the optimal training point for a learner at roughly 85% success: hard enough to learn from, not so hard that the signal is noise. The same band is what Vygotsky's zone of proximal development, Ericsson's edge of ability, and Csikszentmihalyi's flow channel all gesture at. Bloom's mastery learning (advance a unit at >=90% on a fresh check), Bayesian Knowledge Tracing (advance at p(mastery) >= 0.95), and competence-based curriculum learning (Platanios et al. 2019, only attempt tasks within the current competence) are the same rule at different resolutions.
The danger this guards against is specific. An agent that one-shots a large change is maximally helpful to throughput and quietly corrosive to verification: you cannot review what you did not watch get built, and automation bias means you will trust it precisely when it is wrong (Perry et al. 2023). Aviation named the endpoint "children of the magenta": ramp the operator's autonomy faster than their retained understanding and they can no longer hand-fly or override the automation when it misbehaves.
Do not design the whole system up front. Pick the smallest slice that is a real, end-to-end step and stop there. The default rung is about 40 added lines: a change a human can read and explain in one sitting. The bound is the point, not a nuisance: it keeps understanding in pace with output. The guard_scope_ramp.py hook makes this concrete by flagging an increment that jumps past the current rung.
The next increment may be more ambitious only after the prior one's understanding is demonstrated and recorded. The check is sized to blast radius, the advancement gate:
slice has green tests and a recorded tradeoff (what was chosen, what was rejected, why).
crypto): ramp only when the human explains the prior diff unaided. This is the magenta hand-fly check. If they cannot explain it, the rung drops rather than rises.
Recording the demonstration mints a ramp token (touch .imbue/ramp-ok), which the hook consumes to widen the rung one notch. You ramp by proving you understood the last slice, not by writing more. Each notch is appended to the ramp ledger so a reviewer can later audit that the demonstration was real, not rubber-stamped.
Advancing too fast is one failure; never advancing is the other.
shrink the increment, re-scaffold. Do not ramp.
notch.
faster. Drilling a mastered skill is over-practice, the boredom failure that gets spaced-repetition decks abandoned (Cen & Koedinger 2007).
the human will maintain or be accountable for it.
throwaway.
Skip it for a single bounded edit, a trivial reversible change, or generated and vendored code. Forcing a ramp ritual on a typo fix is ceremony, and ceremony trains people to ignore the gate.
imbue:proof-of-work)
| Thought | Reality | |---------|---------| | "I'll just build the whole thing, then review" | You cannot review what you did not watch get built. Start with one slice. | | "Tests pass, so it is understood" | Completion is not understanding. Duolingo streaks prove a cheap signal decouples from skill. | | "I can self-certify I get it" | The producer may not grade its own readiness. Demonstrate it, record it. | | "Bigger increments are faster" | Faster to write, slower to verify, and the verification is the point. | | "The rung is slowing me down" | On work you must own, staying in the 85% band is the fast path to durable skill. |
imbue:assisted-mastery: fades scaffolding as competence grows;this skill ramps challenge. Two directions, one axis.
imbue:proof-of-work: the evidence half of the low-stakes gate.imbue:scope-guard: bounds the branch; this bounds theincrement within it.
leyline:risk-classification: the stakes tier that selects whichgate (evidence vs explanation) applies.
leyline:decision-journal: the durable home for the recordedtradeoff that mints a ramp token.
The empirical basis for the 85% band, the failure modes, and the cross-domain gate design is preserved in research-basis.md.
start rung, not the whole design.
recorded demonstration of the prior increment (a tradeoff entry; for high-stakes paths, the human explaining the diff unaided).
explanation gate, not defaulted silently.
triggered a hold and a smaller next slice, not a ramp.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 15,169 | 12,805 | -16% | 1 | 1 | 0% | 2,702 | 3,627 | +34% | 0 | 0 | — |
case-02 | fail→fail | 13,591 | 11,117 | -18% | 1 | 1 | 0% | 2,166 | 3,442 | +59% | 0 | 0 | — |
case-03 | fail→pass | 16,337 | 10,836 | -34% | 1 | 1 | 0% | 2,723 | 3,355 | +23% | 0 | 0 | — |
case-04 | pass→pass | 10,239 | 2,902 | -72% | 1 | 1 | 0% | 1,483 | 1,970 | +33% | 0 | 0 | — |
case-05 | pass→pass | 9,999 | 4,548 | -55% | 1 | 1 | 0% | 1,483 | 2,264 | +53% | 0 | 0 | — |
case-06 | pass→pass | 12,718 | 7,651 | -40% | 1 | 1 | 0% | 1,964 | 2,627 | +34% | 0 | 0 | — |
case-15 | fail→pass | 13,416 | 8,762 | -35% | 1 | 1 | 0% | 2,082 | 2,794 | +34% | 0 | 0 | — |
case-07 | pass→pass | 11,310 | 3,772 | -67% | 1 | 1 | 0% | 1,602 | 2,189 | +37% | 0 | 0 | — |
case-08 | pass→pass | 15,985 | 4,185 | -74% | 1 | 1 | 0% | 2,324 | 2,207 | -5% | 0 | 0 | — |
case-09 | fail→pass | 12,290 | 5,167 | -58% | 1 | 1 | 0% | 1,950 | 2,426 | +24% | 0 | 0 | — |
case-10 | pass→pass | 10,367 | 4,891 | -53% | 1 | 1 | 0% | 1,588 | 2,325 | +46% | 0 | 0 | — |
case-11 | fail→pass | 9,918 | 4,302 | -57% | 1 | 1 | 0% | 1,641 | 2,240 | +37% | 0 | 0 | — |
case-12 | fail→pass | 12,104 | 6,099 | -50% | 1 | 1 | 0% | 2,049 | 2,594 | +27% | 0 | 0 | — |
case-13 | fail→pass | 13,032 | 9,598 | -26% | 1 | 1 | 0% | 2,301 | 3,204 | +39% | 0 | 0 | — |
case-14 | fail→pass | 15,162 | 8,720 | -42% | 1 | 1 | 0% | 2,405 | 2,966 | +23% | 0 | 0 | — |
case-16 | fail→pass | 10,693 | 3,596 | -66% | 1 | 1 | 0% | 1,727 | 2,119 | +23% | 0 | 0 | — |
case-17 | fail→pass | 11,867 | 2,601 | -78% | 1 | 1 | 0% | 1,841 | 2,015 | +9% | 0 | 0 | — |
case-18 | pass→pass | 6,150 | 3,130 | -49% | 1 | 1 | 0% | 922 | 2,058 | +123% | 0 | 0 | — |
case-19 | pass→pass | 10,564 | 5,421 | -49% | 1 | 1 | 0% | 1,637 | 2,366 | +45% | 0 | 0 | — |
case-20 | pass→pass | 10,924 | 4,592 | -58% | 1 | 1 | 0% | 1,644 | 2,328 | +42% | 0 | 0 | — |
case-21 | pass→pass | 11,985 | 8,006 | -33% | 1 | 1 | 0% | 1,865 | 2,946 | +58% | 0 | 0 | — |
case-22 | fail→pass | 11,188 | 2,809 | -75% | 1 | 1 | 0% | 1,752 | 2,010 | +15% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +50 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.