Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Optionally analyze collections of durable
.claude/skills/boshu2-learn/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | 65% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 42% | 0% |
| case-06 | ✗→✓ | ▲ Improved | -17% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -50% | 0% |
| case-11 | ✗→✓ | ▲ Improved | -64% | 0% |
Learn is an optional, off-path consumer of durable verdict.v2 collections. It may summarize recurring evidence and propose a candidate deterministic check for later human or caller evaluation.
Learn does not run during RPI, validate a subject, alter a verdict, mutate a plan, promote a rule, choose continuation, or mint lifecycle artifacts. Missing Learn output never changes whether a candidate is valid.
When invoked, bind every observation to verdict and finding digests, distinguish repeated objectives from repeated reviews of one objective, disclose the sample size, and stop at advisory evidence.
Overweight failures: a NOT_PROVEN or FAIL verdict carries more teaching value than a PASS, because it names a rule the loop lacked. Harvest kernels from failed lanes first — the canonical example is the mutating-check quarantine in skills/validate/SKILL.md, a durable rule minted from a NOT_PROVEN-then-PASS verdict pair.
Prune for provenance decay: every cited artifact must still resolve — the file exists or the verdict digest is present under .agents/ao/verdicts/. A citation that no longer resolves gets pruned rather than paraphrased, and confidence in a lesson that has not been reproduced since its source decayed goes down, not sideways.
When the caller asks for a durable artifact, write the observations under .agents/scratch/learn/ and return the path; otherwise return them inline. The write is advisory and TTL'd — it is never a source of record, and its absence never changes whether a candidate is valid.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 9,639 | 4,199 | -56% | 1 | 1 | 0% | 958 | 840 | -12% | 0 | 0 | — |
case-02 | fail→fail | 4,235 | 6,546 | +55% | 1 | 1 | 0% | 179 | 733 | +309% | 0 | 0 | — |
case-03 | fail→fail | 13,838 | 13,787 | -0% | 1 | 1 | 0% | 2,262 | 2,498 | +10% | 0 | 0 | — |
case-04 | fail→pass | 5,659 | 6,786 | +20% | 1 | 1 | 0% | 876 | 1,447 | +65% | 0 | 0 | — |
case-05 | fail→pass | 5,018 | 5,107 | +2% | 1 | 1 | 0% | 806 | 1,144 | +42% | 0 | 0 | — |
case-06 | fail→pass | 11,169 | 5,349 | -52% | 1 | 1 | 0% | 1,232 | 1,026 | -17% | 0 | 0 | — |
case-07 | fail→pass | 10,812 | 3,632 | -66% | 1 | 1 | 0% | 1,768 | 883 | -50% | 0 | 0 | — |
case-08 | pass→pass | 13,130 | 6,378 | -51% | 1 | 1 | 0% | 2,145 | 1,393 | -35% | 0 | 0 | — |
case-09 | pass→pass | 4,706 | 3,165 | -33% | 1 | 1 | 0% | 659 | 778 | +18% | 0 | 0 | — |
case-10 | pass→pass | 10,077 | 2,737 | -73% | 1 | 1 | 0% | 1,454 | 764 | -47% | 0 | 0 | — |
case-11 | fail→pass | 11,591 | 1,766 | -85% | 1 | 1 | 0% | 1,769 | 641 | -64% | 0 | 0 | — |
case-12 | fail→fail | 6,013 | 1,737 | -71% | 1 | 1 | 0% | 963 | 578 | -40% | 0 | 0 | — |
case-17 | pass→pass | 11,763 | 2,928 | -75% | 1 | 1 | 0% | 1,757 | 749 | -57% | 0 | 0 | — |
case-13 | fail→fail | 12,661 | 4,130 | -67% | 1 | 1 | 0% | 1,859 | 1,015 | -45% | 0 | 0 | — |
case-14 | pass→pass | 12,711 | 8,067 | -37% | 1 | 1 | 0% | 1,904 | 1,653 | -13% | 0 | 0 | — |
case-15 | fail→pass | 7,824 | 3,448 | -56% | 1 | 1 | 0% | 1,187 | 902 | -24% | 0 | 0 | — |
case-16 | fail→pass | 12,076 | 1,590 | -87% | 1 | 1 | 0% | 1,803 | 566 | -69% | 0 | 0 | — |
case-18 | pass→pass | 13,773 | 3,263 | -76% | 1 | 1 | 0% | 1,967 | 859 | -56% | 0 | 0 | — |
case-19 | fail→fail | 7,330 | 2,912 | -60% | 1 | 1 | 0% | 1,082 | 690 | -36% | 0 | 0 | — |
case-20 | pass→pass | 11,188 | 3,459 | -69% | 1 | 1 | 0% | 1,629 | 849 | -48% | 0 | 0 | — |
case-21 | pass→pass | 10,362 | 3,069 | -70% | 1 | 1 | 0% | 1,592 | 763 | -52% | 0 | 0 | — |
case-22 | pass→pass | 13,189 | 2,684 | -80% | 1 | 1 | 0% | 2,008 | 742 | -63% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 21 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +32 percentage points is the difference between those two pass rates over the 21 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.