Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when a run is finished and the report is written, after the report is written, to record one reusable lesson for the next run in this field. Covers what counts as a lesson worth passing on, what must never be passed on, and how to write it.
.claude/skills/tangxiangru-record-what-you-learned/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 105% | 0% |
| case-08 | ✗→✓ | ▲ Improved | -1% | 0% |
| case-14 | ✗→✓ | ▲ Improved | -26% | 0% |
| case-16 | ✗→✓ | ▲ Improved | -7% | 0% |
| case-17 | ✗→✓ | ▲ Improved | 222% | 0% |
You hit things this run that were in no prompt: an archive that stores its axis in an unexpected order, a reference implementation needing an undocumented flag, a check that caught a mistake you would otherwise have shipped, a step the field treats as obvious and no instruction mentioned.
The next run in this field hits the same thing unless you write it down.
Run this once, at the end:
python3 -c "
import sys; sys.path.insert(0, '<AUTOR_ROOT>')
from src.skill_evolution import record_note
note, problems = record_note(
discipline='<the field, e.g. earth>',
title='<short, routable, what the lesson is about>',
body='''<what you hit, and what to do instead next time>''',
learned_in='<this task id>')
print(problems or 'recorded')
"A good note is one paragraph answering: what surprised you, how it shows up, and what to do instead. Write it for someone competent who has not seen this corpus.
No results. Not your numbers, not the paper's, not "it came out around X". A note travels to a different task, and a finding that travels is contamination — it invites the next run to expect an answer instead of measuring one. The recorder refuses notes containing measured values, and that refusal is not an obstacle to work around.
Nothing you did not hit. A guess about what might help is prose, and prose accumulates until nobody reads the pool. A run that learned nothing transferable records nothing. That is a valid outcome.
Not five. The pool is capped, and a run filing five pushes out four another run earned. Pick the one you most wish you had known at the start.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 21,320 | 20,463 | -4% | 1 | 1 | 0% | 1,972 | 4,033 | +105% | 0 | 0 | — |
case-02 | pass→pass | 9,287 | 11,325 | +22% | 1 | 1 | 0% | 1,515 | 2,137 | +41% | 0 | 0 | — |
case-03 | pass→fail | 11,584 | 9,289 | -20% | 1 | 1 | 0% | 1,956 | 600 | -69% | 0 | 0 | — |
case-04 | fail→fail | 9,554 | 14,124 | +48% | 1 | 1 | 0% | 1,324 | 772 | -42% | 0 | 0 | — |
case-05 | fail→fail | 29,599 | 23,735 | -20% | 1 | 1 | 0% | 1,207 | 760 | -37% | 0 | 0 | — |
case-06 | fail→fail | 23,142 | 6,686 | -71% | 1 | 1 | 0% | 2,904 | 789 | -73% | 0 | 0 | — |
case-07 | fail→fail | 20,872 | 43,134 | +107% | 1 | 1 | 0% | 796 | 955 | +20% | 0 | 0 | — |
case-08 | fail→pass | 13,398 | 9,297 | -31% | 1 | 1 | 0% | 1,926 | 1,900 | -1% | 0 | 0 | — |
case-09 | fail→fail | 12,777 | 23,126 | +81% | 1 | 1 | 0% | 1,898 | 1,034 | -46% | 0 | 0 | — |
case-10 | fail→fail | 24,418 | 21,389 | -12% | 1 | 1 | 0% | 4,178 | 943 | -77% | 0 | 0 | — |
case-11 | pass→pass | 21,708 | 7,422 | -66% | 1 | 1 | 0% | 1,661 | 1,670 | +1% | 0 | 0 | — |
case-12 | fail→fail | 20,060 | 15,203 | -24% | 1 | 1 | 0% | 3,510 | 875 | -75% | 0 | 0 | — |
case-13 | fail→fail | 33,736 | 9,913 | -71% | 1 | 1 | 0% | 498 | 871 | +75% | 0 | 0 | — |
case-14 | fail→pass | 11,037 | 8,764 | -21% | 1 | 1 | 0% | 1,936 | 1,434 | -26% | 0 | 0 | — |
case-15 | fail→fail | 9,351 | 8,758 | -6% | 1 | 1 | 0% | 1,412 | 764 | -46% | 0 | 0 | — |
case-16 | fail→pass | 13,805 | 5,787 | -58% | 1 | 1 | 0% | 1,328 | 1,231 | -7% | 0 | 0 | — |
case-17 | fail→pass | 5,070 | 19,741 | +289% | 1 | 1 | 0% | 796 | 2,566 | +222% | 0 | 0 | — |
case-18 | fail→fail | 16,643 | 13,010 | -22% | 1 | 1 | 0% | 2,648 | 822 | -69% | 0 | 0 | — |
case-19 | fail→fail | 7,785 | 25,690 | +230% | 1 | 1 | 0% | 1,342 | 5,219 | +289% | 0 | 0 | — |
case-20 | fail→pass | 7,378 | 5,914 | -20% | 1 | 1 | 0% | 1,069 | 1,171 | +10% | 0 | 0 | — |
case-21 | pass→pass | 9,253 | 4,099 | -56% | 1 | 1 | 0% | 1,383 | 842 | -39% | 0 | 0 | — |
case-22 | fail→fail | 15,119 | 21,120 | +40% | 1 | 1 | 0% | 974 | 792 | -19% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 10 counted toward the lift figure. The other 12 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +23 percentage points is the difference between those two pass rates over the 10 comparable cases. 3 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.