Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Run a portable sequential plan-work-verify-review-compound lifecycle. Use optional generic workers only when Amp makes them available.
.claude/skills/oliver-kriska-phx-full/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-16 | ✗→✓ | ▲ Improved | 42% | 0% |
| case-19 | ✗→✓ | ▲ Improved | 30% | 0% |
| case-20 | ✗→✓ | ▲ Improved | 40% | 0% |
| case-21 | ✗→✓ | ▲ Improved | -38% | 0% |
| case-13 | ✓→✗ | ▼ Worse | -34% | 0% |
Run the portable lifecycle: discover → plan → work → verify → read-only review → compound. The filesystem is the state machine; no task API or named orchestrator is required.
text/skill:phx-full Add user authentication with magic links /skill:phx-full Background email jobs --max-cycles 5 --max-retries 2
If input is an existing .claude/plans/*/plan.md, do not re-plan. Ask for the native phx-work workflow or execute its portable behavior in this session. Defaults are --max-cycles 10, --max-retries 3, and --max-blockers 5.
Tidewave evidence. Tidewave is optional; local files, logs, and mix commands are the complete fallback. Record complexity and proposed depth, then wait for the user's plan/implementation gate. Never auto-select a path that bypasses it.
phx-plan skill when available, orexecute its portable research checklist and artifact format in this session. Require .claude/plans/{slug}/plan.md. Present it and wait for approval before implementation unless the user already explicitly authorized the full run.
The full-run limits override any baseline workflow retry defaults. Before every attempt persist cycle, task retry, and blocker counters; if the next attempt exceeds a limit, do not run it. --max-retries N means at most N retries after the initial attempt (N+1 total attempts for that task). Mark [BLOCKED] and stop at --max-blockers.
mix format --check-formatted, compile with warnings aserrors, focused tests during work, and the full relevant suite at this gate. A failed gate appends FAIL and returns to WORKING only within the cycle limit.
phx-review, or perform the same read-only,changed-file review sequentially. Generic workers are optional. Review never edits. Findings or failures become plan tasks and return to WORKING.
invoke phx-compound. Inline contract: write a solution artifact under .claude/solutions/ only when the run produced a non-obvious, reusable learning, including problem, root cause, solution, and verification. Otherwise append COMPOUNDING SKIPPED: no reusable learning to progress. Never edit CLAUDE.md.
Track INITIALIZING → DISCOVERING → PLANNING → WORKING → VERIFYING → REVIEWING → COMPOUNDING → COMPLETED, with BLOCKED reachable from every phase. A cycle is one WORKING → VERIFYING → REVIEWING pass; increment and persist it before entering VERIFYING. At --max-cycles, do not begin another pass: stop INCOMPLETE with remaining tasks, failed evidence, and a concrete resume command for this runtime.
progress.md is the sole state authority. It is append-only: never overwrite or maintain a competing authoritative current-state record. Every event has monotonic seq, phase_visit, phase, cycle, task, task_attempt, cumulative blockers, outcome, and an evidence or artifact path. On resume, validate the last valid event against evidence, plan checkboxes, artifacts, and git state, then enter only its legal successor. Any WORKING edit after a VERIFYING or REVIEWING pass invalidates both passes; the next legal phase is VERIFYING.
Completion requires all required plan tasks checked, no unresolved [BLOCKED], the latest VERIFYING PASS after the last edit, the latest accepted REVIEWING after that verify, and COMPOUNDING passed or explicitly skipped.
references/execution-steps.md — portable phase gates and outputsreferences/example-run.md — example lifecyclereferences/safety-recovery.md — resume and blocker recoveryreferences/cycle-patterns.md — bounded cycle patterns| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-02 | fail→fail | 4,306 | 4,748 | +10% | 1 | 1 | 0% | 213 | 1,435 | +574% | 0 | 0 | — |
case-01 | fail→fail | 5,324 | 4,868 | -9% | 1 | 1 | 0% | 233 | 1,392 | +497% | 0 | 0 | — |
case-03 | fail→fail | 4,513 | 6,846 | +52% | 1 | 1 | 0% | 200 | 1,408 | +604% | 0 | 0 | — |
case-04 | fail→fail | 2,863 | 5,355 | +87% | 1 | 1 | 0% | 422 | 1,450 | +244% | 0 | 0 | — |
case-05 | fail→fail | 4,782 | 5,431 | +14% | 1 | 1 | 0% | 305 | 1,340 | +339% | 0 | 0 | — |
case-06 | pass→pass | 17,379 | 15,864 | -9% | 1 | 1 | 0% | 3,008 | 3,849 | +28% | 0 | 0 | — |
case-07 | fail→fail | 16,829 | 7,608 | -55% | 1 | 1 | 0% | 2,877 | 1,714 | -40% | 0 | 0 | — |
case-08 | fail→fail | 22,996 | 4,580 | -80% | 1 | 1 | 0% | 4,379 | 1,288 | -71% | 0 | 0 | — |
case-09 | fail→fail | 4,622 | 7,273 | +57% | 1 | 1 | 0% | 183 | 1,400 | +665% | 0 | 0 | — |
case-10 | fail→fail | 4,106 | 5,302 | +29% | 1 | 1 | 0% | 199 | 1,413 | +610% | 0 | 0 | — |
case-11 | fail→fail | 6,181 | 6,401 | +4% | 1 | 1 | 0% | 866 | 1,451 | +68% | 0 | 0 | — |
case-12 | pass→pass | 5,855 | 7,454 | +27% | 1 | 1 | 0% | 790 | 1,777 | +125% | 0 | 0 | — |
case-18 | fail→fail | 6,595 | 5,773 | -12% | 1 | 1 | 0% | 1,077 | 1,479 | +37% | 0 | 0 | — |
case-13 | pass→fail | 12,759 | 4,933 | -61% | 1 | 1 | 0% | 2,071 | 1,359 | -34% | 0 | 0 | — |
case-14 | fail→fail | 16,232 | 5,726 | -65% | 1 | 1 | 0% | 2,213 | 1,505 | -32% | 0 | 0 | — |
case-15 | fail→fail | 7,942 | 6,304 | -21% | 1 | 1 | 0% | 726 | 1,563 | +115% | 0 | 0 | — |
case-16 | fail→pass | 8,192 | 4,310 | -47% | 1 | 1 | 0% | 1,310 | 1,866 | +42% | 0 | 0 | — |
case-17 | fail→fail | 5,543 | 5,934 | +7% | 1 | 1 | 0% | 286 | 1,512 | +429% | 0 | 0 | — |
case-19 | fail→pass | 8,660 | 3,434 | -60% | 1 | 1 | 0% | 1,287 | 1,671 | +30% | 0 | 0 | — |
case-20 | fail→pass | 11,075 | 5,296 | -52% | 1 | 1 | 0% | 1,431 | 2,009 | +40% | 0 | 0 | — |
case-21 | fail→pass | 16,105 | 3,290 | -80% | 1 | 1 | 0% | 2,518 | 1,570 | -38% | 0 | 0 | — |
case-22 | pass→fail | 12,364 | 3,577 | -71% | 1 | 1 | 0% | 2,253 | 1,377 | -39% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 8 counted toward the lift figure. The other 14 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +9 percentage points is the difference between those two pass rates over the 8 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.