Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Self-improving loop for plugin skills. Reads program.md, proposes one mutation per iteration, evaluates against deterministic scorer, keeps improvements via git, reverts failures. Targets weakest skill+dimension. Use with /loop for overnight runs.
.claude/skills/majiayu000-lab-autoresearch/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-05 | ✗→✓ | ▲ Improved | -11% | 0% |
| case-06 | ✗→✓ | ▲ Improved | -13% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 34% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 20% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 5% | 0% |
Iteratively improve plugin skills via the autoresearch pattern: propose one mutation -> eval -> keep/revert -> repeat.
/lab:autoresearch # Targeted: attack weakest skill+dimension
/lab:autoresearch --skill review # Focus on one skill
/lab:autoresearch --strategy sweep # Process all skills alphabetically
/lab:autoresearch --dry-run # Show what would change, don't commitFor overnight runs:
/loop 5m /lab:autoresearch --strategy sweep --max-iterations 200keep or revert command (never skip)All eval/git/journal operations go through ONE script. Do NOT run these manually.
bash# Find the weakest skill+dimension python3 lab/autoresearch/scripts/run-iteration.py target --strategy targeted # Score a skill (before mutation, to get baseline) python3 lab/autoresearch/scripts/run-iteration.py score <skill-name> # After mutation: score + checks + compare → verdict (KEEP or REVERT) python3 lab/autoresearch/scripts/run-iteration.py eval <skill-name> # Act on verdict: python3 lab/autoresearch/scripts/run-iteration.py keep <skill> <dim> <old> <new> \ --desc "what changed" --asi '{"hypothesis": "why", "mechanism": "how"}' python3 lab/autoresearch/scripts/run-iteration.py revert <skill> <dim> <old> <new> \ --desc "what was attempted" --asi '{"hypothesis": "why", "regression": "what broke", "avoid": "do not retry this"}' # Check overall progress python3 lab/autoresearch/scripts/run-iteration.py status
lab/autoresearch/program.md (goals, mutable surface, rules)lab/autoresearch/ideas.md if it exists (deferred optimizations)python3 lab/autoresearch/scripts/run-iteration.py statusRun: python3 lab/autoresearch/scripts/run-iteration.py target --strategy targeted
Parse the JSON: skill, dimension, failing_checks. If all_perfect → STOP.
lab/eval/evals/{skill}.jsonideas.md for deferred ideas about this skill${CLAUDE_SKILL_DIR}/references/mutation-strategies.mdpython3 lab/autoresearch/scripts/run-iteration.py eval <skill-name>verdict fieldIf verdict is KEEP:
bashpython3 lab/autoresearch/scripts/run-iteration.py keep <skill> <dim> <old> <new> \ --desc "..." --asi '{"hypothesis": "...", "mechanism": "..."}'
If verdict is REVERT:
bashpython3 lab/autoresearch/scripts/run-iteration.py revert <skill> <dim> <old> <new> \ --desc "..." --asi '{"hypothesis": "...", "regression": "...", "avoid": "..."}'
If during analysis you discovered a promising optimization you can't act on now:
lab/autoresearch/ideas.md as a bullet${CLAUDE_SKILL_DIR}/references/mutation-strategies.md — mutation type catalog${CLAUDE_SKILL_DIR}/references/state-management.md — git protocol, journalinglab/autoresearch/program.md — research agenda (read every iteration)Other measured skills in the registry, with their headline benchmark lift.