Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Guide the agent to recall, remember, and route durable learning into Memory, Skills, Scheduled Tasks, or Tape.
.claude/skills/thinkinaixyz-memory-management/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-10 | ✗→✓ | ▲ Improved | -31% | 0% |
| case-23 | ✗→✓ | ▲ Improved | 2% | 0% |
| case-20 | ✓→✗ | ▼ Worse | 199% | 0% |
| case-09 | ✓→✓ | = Same ✓ | 84% | 0% |
| case-11 | ✓→✓ | = Same ✓ | -36% | 0% |
Use this skill when a task may produce durable learning or when the user asks you to recall, remember, continue earlier work, preserve an exact statement, capture a reusable procedure, or handle a recurring need.
Rely on automatic memory injection for ordinary context. Use memory_recall when the user refers to previous work with cues such as again, last time, before, continue, same project, remember, or asks what you already know.
Use tape_search and then tape_context when the user needs source evidence, exact wording, logs, command output, file snippets, or why a prior decision was made. Memory is a durable conclusion layer, not the raw transcript.
Use memory_remember only for durable conclusions that should change future behavior. Choose the most specific category:
user_preference: stable user preferences, constraints, communication style, environment choices.project_fact: durable project conventions, architecture entry points, commands, dependencies, paths, or operational constraints.task_outcome: completed, blocked, or deliberately deferred task results. Include status, outcome, and blocker in prose when relevant.heuristic: reusable troubleshooting strategy, workflow, decision rule, or engineering lesson.anti_pattern: repeated mistake, unsafe approach, brittle pattern, stale assumption, or thing to avoid.Do not remember raw tool results, bash output, grep output, file contents, transient mechanics, one-off failures, secrets, credentials, hidden reasoning, or anything only useful for the current turn.
Store exact wording only when the user explicitly asks you to remember a sentence or phrase verbatim. In that case, keep the requested text intact and make the surrounding content minimal.
Automatic extraction is different: it should normalize durable facts into concise memory content, deduplicate related entries, and avoid preserving raw transcript text.
When the useful learning is a reusable multi-step procedure, prefer drafting a skill with skill_manage instead of stuffing the full procedure into Memory. Memory may keep a short pointer or heuristic, but the repeatable workflow belongs in a Skill.
Use skill_manage for draft skills only. Do not modify installed skills unless the user explicitly asks through the supported review flow.
When the user asks for a periodic, low-frequency, or future recurring action, suggest creating a Scheduled Task in settings. Memory does not wake the agent, schedule future work, or create automation side effects.
Before finishing a non-trivial task, check whether there is one durable lesson to save:
skill_manage or a recurring need for Scheduled Tasks rather than Memory?Remember only the smallest durable conclusion. Leave raw process in Tape.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 3,793 | 7,722 | +104% | 1 | 1 | 0% | 609 | 1,533 | +152% | 0 | 0 | — |
case-02 | fail→fail | 13,999 | 37,894 | +171% | 1 | 1 | 0% | 2,567 | 4,357 | +70% | 0 | 0 | — |
case-03 | fail→fail | 3,372 | 6,864 | +104% | 1 | 1 | 0% | 541 | 1,048 | +94% | 0 | 0 | — |
case-04 | fail→fail | 2,665 | 5,297 | +99% | 1 | 1 | 0% | 483 | 961 | +99% | 0 | 0 | — |
case-05 | fail→fail | 3,349 | 4,942 | +48% | 1 | 1 | 0% | 578 | 983 | +70% | 0 | 0 | — |
case-06 | fail→fail | 13,494 | 10,890 | -19% | 1 | 1 | 0% | 2,290 | 1,768 | -23% | 0 | 0 | — |
case-07 | fail→fail | 12,054 | 4,968 | -59% | 1 | 1 | 0% | 2,008 | 952 | -53% | 0 | 0 | — |
case-08 | fail→fail | 1,947 | 5,353 | +175% | 1 | 1 | 0% | 263 | 892 | +239% | 0 | 0 | — |
case-09 | pass→pass | 6,133 | 4,865 | -21% | 1 | 1 | 0% | 648 | 1,193 | +84% | 0 | 0 | — |
case-10 | fail→pass | 11,248 | 5,049 | -55% | 1 | 1 | 0% | 2,182 | 1,497 | -31% | 0 | 0 | — |
case-11 | pass→pass | 10,886 | 2,750 | -75% | 1 | 1 | 0% | 1,748 | 1,116 | -36% | 0 | 0 | — |
case-17 | fail→fail | 4,742 | 7,135 | +50% | 1 | 1 | 0% | 797 | 1,336 | +68% | 0 | 0 | — |
case-12 | pass→pass | 8,294 | 3,293 | -60% | 1 | 1 | 0% | 1,388 | 1,198 | -14% | 0 | 0 | — |
case-13 | fail→fail | 9,132 | 6,738 | -26% | 1 | 1 | 0% | 1,488 | 1,200 | -19% | 0 | 0 | — |
case-14 | fail→fail | 3,819 | 3,798 | -1% | 1 | 1 | 0% | 640 | 885 | +38% | 0 | 0 | — |
case-15 | fail→fail | 3,187 | 4,157 | +30% | 1 | 1 | 0% | 490 | 845 | +72% | 0 | 0 | — |
case-16 | fail→fail | 3,928 | 4,613 | +17% | 1 | 1 | 0% | 678 | 920 | +36% | 0 | 0 | — |
case-18 | fail→fail | 2,321 | 5,586 | +141% | 1 | 1 | 0% | 324 | 974 | +201% | 0 | 0 | — |
case-19 | fail→fail | 10,427 | 5,684 | -45% | 1 | 1 | 0% | 1,990 | 1,017 | -49% | 0 | 0 | — |
case-20 | pass→fail | 2,231 | 5,693 | +155% | 1 | 1 | 0% | 359 | 1,073 | +199% | 0 | 0 | — |
case-21 | fail→fail | 4,589 | 10,628 | +132% | 1 | 1 | 0% | 942 | 1,533 | +63% | 0 | 0 | — |
case-22 | fail→fail | 4,486 | 3,993 | -11% | 1 | 1 | 0% | 712 | 878 | +23% | 0 | 0 | — |
case-23 | fail→pass | 7,108 | 8,384 | +18% | 1 | 1 | 0% | 1,187 | 1,207 | +2% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 6 counted toward the lift figure. The other 17 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +4 percentage points is the difference between those two pass rates over the 6 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.