Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Plan features spanning multiple domains: billing (Stripe), auth (RBAC), real-time (Presence), webhooks, jobs (Oban). Use when designing interconnected systems or converting review findings into tasks.
.claude/skills/oliver-kriska-phx-plan/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-15 | ✗→✓ | ▲ Improved | 85% | 0% |
| case-05 | ✓→✗ | ▼ Worse | -38% | 0% |
| case-06 | ✓→✗ | ▼ Worse | -41% | 0% |
| case-16 | ✓→✗ | ▼ Worse | -30% | 0% |
| case-19 | ✓→✗ | ▼ Worse | -44% | 0% |
Plan a feature by researching the relevant Elixir/Phoenix concerns, then output a structured plan with checkboxes.
[ecto], [liveview], [oban] task routingmix compile/format/credo/test verification/skill:phx-plan Add user avatars with S3 upload
/skill:phx-plan .claude/plans/notifications/reviews/notifications-review.md
/skill:phx-plan Implement notifications --depth deep
/skill:phx-plan .claude/plans/auth/plan.md --existing--depth quick|standard|deep = Planning depth (auto-detected)--existing = Enhance an existing plan with deeper researchinterview.md (skip clarification), clear description, or vague
brainstorm interview.md exists with Status: COMPLETE)
.claude/plans/{slug}/scratchpad.md with a concern-track checklist
configured and exposed; otherwise inspect source, routes, schemas, and tests
subagents may run independent tracks in parallel, but they are optional. Without them, run every selected track sequentially in this session and save evidence under .claude/plans/{slug}/research/
marking each selected track [x] only after its evidence is captured. NEVER write the plan while any selected track remains unchecked
Reuse .claude/plans/{slug}/scratchpad.md for decisions and dead-ends
When planning from review: Every finding must appear in the plan — either as a task OR explicitly deferred by the user.
See references/planning-workflow.md for detailed step-by-step.
Enhance an existing plan without relying on named agents:
.claude/plans/{slug}/scratchpad.mdare optional only for independent tracks and must write evidence under .claude/plans/{slug}/research/
evidence, independent of whether workers were used
input is a review file or /skill:phx-investigate output, the findings ARE the research. Do NOT spawn agents to re-discover what the review already found. Convert findings directly to plan tasks. (Confirmed: 56-session analysis showed same findings discovered 3-4x across review→investigate→plan phases, wasting ~96K tokens)
text/skill:phx-plan {feature} <-- YOU ARE HERE | /skill:phx-plan --existing (optional enhancement) | ASK USER -> /skill:phx-work .claude/plans/{feature}/plan.md | /skill:phx-review → /skill:phx-compound
.claude/plans/{slug}/plan.md.claude/plans/{slug}/research/ can be deleted afterSTOP. Do NOT proceed to implementation.
After writing .claude/plans/{slug}/plan.md:
/skill:phx-brief — interactive walkthrough)When user selects "Start in fresh session", print:
1. Run `/new` to start a fresh session
2. Then run one of:
/skill:phx-work .claude/plans/{slug}/plan.md
/skill:phx-full .claude/plans/{slug}/plan.md (includes review + compound)This is Iron Law #1. Violating it wastes user context.
references/planning-workflow.md — Detailed step-by-stepreferences/plan-template.mdreferences/complexity-detail.mdreferences/example-plan.mdreferences/agent-selection.mdreferences/breadboarding.md| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 17,664 | 3,967 | -78% | 1 | 1 | 0% | 3,202 | 1,829 | -43% | 0 | 0 | — |
case-02 | fail→fail | 30,840 | 5,972 | -81% | 1 | 1 | 0% | 4,796 | 1,802 | -62% | 0 | 0 | — |
case-03 | fail→fail | 26,345 | 6,103 | -77% | 1 | 1 | 0% | 4,942 | 1,827 | -63% | 0 | 0 | — |
case-04 | fail→fail | 10,555 | 4,573 | -57% | 1 | 1 | 0% | 1,961 | 1,776 | -9% | 0 | 0 | — |
case-05 | pass→fail | 13,804 | 4,353 | -68% | 1 | 1 | 0% | 2,746 | 1,697 | -38% | 0 | 0 | — |
case-06 | pass→fail | 16,568 | 5,440 | -67% | 1 | 1 | 0% | 2,800 | 1,648 | -41% | 0 | 0 | — |
case-07 | fail→fail | 15,765 | 3,705 | -76% | 1 | 1 | 0% | 2,864 | 1,711 | -40% | 0 | 0 | — |
case-08 | fail→fail | 20,268 | 4,997 | -75% | 1 | 1 | 0% | 3,211 | 1,708 | -47% | 0 | 0 | — |
case-09 | fail→fail | 19,466 | 6,287 | -68% | 1 | 1 | 0% | 3,310 | 1,810 | -45% | 0 | 0 | — |
case-10 | fail→fail | 16,738 | 3,989 | -76% | 1 | 1 | 0% | 2,800 | 1,638 | -42% | 0 | 0 | — |
case-11 | fail→fail | 19,179 | 5,673 | -70% | 1 | 1 | 0% | 3,599 | 1,703 | -53% | 0 | 0 | — |
case-12 | fail→fail | 20,950 | 4,190 | -80% | 1 | 1 | 0% | 4,056 | 1,763 | -57% | 0 | 0 | — |
case-13 | fail→fail | 22,373 | 3,510 | -84% | 1 | 1 | 0% | 4,125 | 1,663 | -60% | 0 | 0 | — |
case-14 | fail→fail | 20,031 | 5,420 | -73% | 1 | 1 | 0% | 3,892 | 1,740 | -55% | 0 | 0 | — |
case-15 | fail→pass | 7,430 | 2,973 | -60% | 1 | 1 | 0% | 1,052 | 1,948 | +85% | 0 | 0 | — |
case-16 | pass→fail | 15,078 | 3,037 | -80% | 1 | 1 | 0% | 2,666 | 1,879 | -30% | 0 | 0 | — |
case-17 | fail→fail | 16,795 | 5,665 | -66% | 1 | 1 | 0% | 2,750 | 1,792 | -35% | 0 | 0 | — |
case-18 | fail→fail | 30,322 | 4,044 | -87% | 1 | 1 | 0% | 5,474 | 1,653 | -70% | 0 | 0 | — |
case-19 | pass→fail | 17,306 | 5,065 | -71% | 1 | 1 | 0% | 3,028 | 1,694 | -44% | 0 | 0 | — |
case-20 | fail→fail | 2,383 | 5,504 | +131% | 1 | 1 | 0% | 296 | 1,828 | +518% | 0 | 0 | — |
case-21 | pass→fail | 14,886 | 4,112 | -72% | 1 | 1 | 0% | 2,771 | 1,652 | -40% | 0 | 0 | — |
case-22 | pass→fail | 15,833 | 3,470 | -78% | 1 | 1 | 0% | 2,854 | 1,676 | -41% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 2 counted toward the lift figure. The other 20 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of -23 percentage points is the difference between those two pass rates over the 2 comparable cases. 12 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.