Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when you want Codex to review its own recent history (last N days or specific period) and improve its behavior. Produces minimal, high-signal updates to AGENTS.md and tiny reusable skills. The goal is long-term fluency — Codex gradually becomes better at your specific style, constraints, and workflows.
.claude/skills/majiayu000-codex-retrospective/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 17% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 12% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 24% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 144% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 56% | 0% |
A structured self-improvement loop for Codex.
This skill turns Codex from a one-off collaborator into a system that gets meaningfully better at working with you over time by systematically eating its own usage history.
Most people improve their agent usage by manually maintaining AGENTS.md or skills when they notice problems. This skill makes that process deliberate, regular, and high-leverage.
It is directly inspired by strong practices from heavy users (especially Greg Brockman's emphasis on updating the "constitution" from real failures and friction), but executed as a repeatable arsenal-style workflow.
Minimal effective change, grounded in evidence from actual history.
Every output must be:
codex-fluent reports that Codex is constantly re-asking for the same context or preferencesYou specify the time window or focus area:
Before proposing any AGENTS.md update or tiny skill, Codex must build a small evidence inventory:
Every proposed update must map back to at least one concrete evidence handle.
Codex is instructed to look for:
The skill forces Codex to produce output in this order:
You review. The skill then helps you apply the minimal changes cleanly (never blindly overwriting large sections of AGENTS.md).
These two skills are designed to be used together:
codex-retrospective finds behavioral and knowledge improvements (better defaults, new rules, extracted skills).codex-fluent finds state and context hygiene improvements (session bloat, missing handoffs, archive opportunities).A good monthly ritual for serious users:
codex-retrospective on the last 30 days.codex-fluent diagnosis.unusually high impact. Prefer "no change" over a weak constitution update.
date range, file, PR, or repeated correction pattern that can be checked.
paragraphs, it probably belongs in a tiny skill or a project doc instead.
but it must not silently rewrite AGENTS.md or existing skills.
references/retrospective-prompt.md — The core prompt template used to drive Codex's self-analysisreferences/agents-md-update-rules.md — Strict rules for what kind of changes are acceptablereferences/minimal-skill-criteria.md — What qualifies as a "tiny useful skill" worth extractingreferences/examples/ — Real (sanitized) retrospective outputs and the resulting AGENTS.md diffsAfter 4–8 weeks of regular use:
Start with a focused 7- or 14-day retrospective on a project where you've felt the most friction recently. The pattern will become natural quickly.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 21,453 | 17,236 | -20% | 1 | 1 | 0% | 3,368 | 3,957 | +17% | 0 | 0 | — |
case-02 | fail→pass | 20,729 | 15,885 | -23% | 1 | 1 | 0% | 3,229 | 3,603 | +12% | 0 | 0 | — |
case-03 | fail→pass | 19,717 | 18,166 | -8% | 1 | 1 | 0% | 3,089 | 3,815 | +24% | 0 | 0 | — |
case-04 | fail→pass | 5,936 | 5,740 | -3% | 1 | 1 | 0% | 943 | 2,303 | +144% | 0 | 0 | — |
case-05 | pass→pass | 10,674 | 14,254 | +34% | 1 | 1 | 0% | 1,834 | 3,499 | +91% | 0 | 0 | — |
case-06 | fail→pass | 15,486 | 15,133 | -2% | 1 | 1 | 0% | 2,496 | 3,906 | +56% | 0 | 0 | — |
case-07 | fail→pass | 25,242 | 14,429 | -43% | 1 | 1 | 0% | 3,878 | 3,695 | -5% | 0 | 0 | — |
case-08 | fail→fail | 5,626 | 17,123 | +204% | 1 | 1 | 0% | 825 | 4,245 | +415% | 0 | 0 | — |
case-09 | pass→pass | 14,045 | 21,708 | +55% | 1 | 1 | 0% | 2,182 | 4,203 | +93% | 0 | 0 | — |
case-10 | fail→fail | 15,752 | 8,850 | -44% | 1 | 1 | 0% | 2,423 | 2,710 | +12% | 0 | 0 | — |
case-11 | fail→pass | 20,586 | 16,113 | -22% | 1 | 1 | 0% | 3,235 | 4,309 | +33% | 0 | 0 | — |
case-12 | fail→pass | 10,786 | 6,279 | -42% | 1 | 1 | 0% | 1,617 | 2,335 | +44% | 0 | 0 | — |
case-13 | fail→pass | 23,657 | 15,037 | -36% | 1 | 1 | 0% | 3,636 | 3,893 | +7% | 0 | 0 | — |
case-14 | fail→pass | 17,661 | 12,458 | -29% | 1 | 1 | 0% | 2,706 | 3,380 | +25% | 0 | 0 | — |
case-15 | fail→pass | 15,669 | 12,995 | -17% | 1 | 1 | 0% | 2,490 | 3,447 | +38% | 0 | 0 | — |
case-16 | pass→pass | 19,892 | 27,117 | +36% | 1 | 1 | 0% | 2,963 | 4,423 | +49% | 0 | 0 | — |
case-17 | fail→pass | 15,439 | 12,355 | -20% | 1 | 1 | 0% | 2,285 | 3,397 | +49% | 0 | 0 | — |
case-18 | fail→pass | 14,035 | 16,970 | +21% | 1 | 1 | 0% | 1,846 | 4,179 | +126% | 0 | 0 | — |
case-19 | pass→pass | 8,895 | 11,121 | +25% | 1 | 1 | 0% | 1,259 | 2,979 | +137% | 0 | 0 | — |
case-20 | pass→pass | 17,139 | 18,553 | +8% | 1 | 1 | 0% | 3,175 | 4,818 | +52% | 0 | 0 | — |
case-21 | pass→pass | 16,575 | 9,753 | -41% | 1 | 1 | 0% | 2,624 | 2,894 | +10% | 0 | 0 | — |
case-22 | fail→pass | 33,440 | 17,023 | -49% | 1 | 1 | 0% | 5,275 | 3,992 | -24% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +64 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.