Install any skill in seconds. Free to start, no credit card required.
Get Started Free →This skill should be used only when the user explicitly asks to use `$ralph-specum-tasks`, or explicitly asks Ralph Specum in Codex to run the tasks phase.
.claude/skills/tzachbon-ralph-specum-tasks/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-12 | ✗→✓ | ▲ Improved | -62% | 0% |
| case-14 | ✗→✓ | ▲ Improved | 64% | 0% |
| case-16 | ✗→✓ | ▲ Improved | -23% | 0% |
| case-20 | ✗→✓ | ▲ Improved | -26% | 0% |
| case-05 | ✓→✗ | ▼ Worse | -68% | 0% |
You are a coordinator, not a task planner -- delegate ALL work to a task-planner sub-agent.
.current-specrequirements.md and design.mdrequirements.md and design.md. Read research.md when present, .progress.md, and current state.awaitingApproval: false before generation.granularity from state. Allow --tasks-size fine|coarse to override it. In quick mode, default unset granularity to fine.task-planner sub-agent. Pass requirements, design, research, and interview context. The sub-agent writes tasks.md. Do NOT write tasks.md yourself.phase: "tasks"awaitingApproval: true (or false when --quick is active)taskIndex: first incomplete or totalTaskstotalTasks: counted tasks.progress.md with the phase breakdown, next milestone, blockers, next step, chosen granularity, and verification strategy.--quick: STOP HERE. Display the walkthrough summary and approval prompt. Do NOT continue to implementation. Wait for the user to explicitly approve and request the next phase.--quick: Review quickly, then continue directly into implementation.Use atomic tasks with exact file targets, explicit success criteria, verification commands, and commit messages. Preserve POC-first ordering. Support [P] markers for safe parallel work, [VERIFY] checkpoints, and VE tasks when end-to-end verification is part of the plan.
tasks.md, name tasks.md and summarize the task plan briefly.approve current artifactrequest changescontinue to implementationcontinue to implementation as approval of tasks.md.| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 11,652 | 7,919 | -32% | 1 | 1 | 0% | 1,966 | 1,090 | -45% | 0 | 0 | — |
case-02 | fail→fail | 2,330 | 5,419 | +133% | 1 | 1 | 0% | 302 | 719 | +138% | 0 | 0 | — |
case-03 | fail→fail | 7,285 | 5,512 | -24% | 1 | 1 | 0% | 1,082 | 983 | -9% | 0 | 0 | — |
case-04 | pass→pass | 20,653 | 16,798 | -19% | 1 | 1 | 0% | 3,403 | 3,234 | -5% | 0 | 0 | — |
case-05 | pass→fail | 13,286 | 4,921 | -63% | 1 | 1 | 0% | 2,670 | 863 | -68% | 0 | 0 | — |
case-06 | pass→fail | 22,482 | 4,250 | -81% | 1 | 1 | 0% | 3,497 | 820 | -77% | 0 | 0 | — |
case-07 | pass→fail | 7,140 | 16,918 | +137% | 1 | 1 | 0% | 1,086 | 3,050 | +181% | 0 | 0 | — |
case-08 | fail→fail | 4,288 | 6,485 | +51% | 1 | 1 | 0% | 576 | 874 | +52% | 0 | 0 | — |
case-09 | fail→fail | 5,589 | 12,147 | +117% | 1 | 1 | 0% | 852 | 1,019 | +20% | 0 | 0 | — |
case-10 | fail→fail | 6,356 | 7,516 | +18% | 1 | 1 | 0% | 825 | 1,011 | +23% | 0 | 0 | — |
case-11 | pass→fail | 3,359 | 2,674 | -20% | 1 | 1 | 0% | 421 | 784 | +86% | 0 | 0 | — |
case-12 | fail→pass | 19,060 | 4,163 | -78% | 1 | 1 | 0% | 3,174 | 1,219 | -62% | 0 | 0 | — |
case-13 | fail→fail | 8,318 | 6,269 | -25% | 1 | 1 | 0% | 1,285 | 896 | -30% | 0 | 0 | — |
case-14 | fail→pass | 4,821 | 4,847 | +1% | 1 | 1 | 0% | 832 | 1,365 | +64% | 0 | 0 | — |
case-15 | fail→fail | 10,166 | 6,810 | -33% | 1 | 1 | 0% | 1,878 | 904 | -52% | 0 | 0 | — |
case-16 | fail→pass | 7,727 | 2,279 | -71% | 1 | 1 | 0% | 1,194 | 914 | -23% | 0 | 0 | — |
case-17 | pass→fail | 6,989 | 6,697 | -4% | 1 | 1 | 0% | 990 | 908 | -8% | 0 | 0 | — |
case-18 | fail→fail | 12,068 | 6,255 | -48% | 1 | 1 | 0% | 1,956 | 1,028 | -47% | 0 | 0 | — |
case-19 | fail→fail | 17,638 | 6,665 | -62% | 1 | 1 | 0% | 2,852 | 791 | -72% | 0 | 0 | — |
case-20 | fail→pass | 15,320 | 7,148 | -53% | 1 | 1 | 0% | 2,281 | 1,699 | -26% | 0 | 0 | — |
case-21 | pass→fail | 8,588 | 4,523 | -47% | 1 | 1 | 0% | 1,163 | 834 | -28% | 0 | 0 | — |
case-22 | fail→fail | 17,850 | 6,222 | -65% | 1 | 1 | 0% | 2,901 | 903 | -69% | 0 | 0 | — |
case-23 | pass→pass | 13,484 | 7,397 | -45% | 1 | 1 | 0% | 2,112 | 1,744 | -17% | 0 | 0 | — |
case-24 | fail→fail | 8,333 | 2,654 | -68% | 1 | 1 | 0% | 1,348 | 999 | -26% | 0 | 0 | — |
case-25 | pass→pass | 7,326 | 2,799 | -62% | 1 | 1 | 0% | 1,088 | 1,022 | -6% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 25 cases were attempted, and 10 counted toward the lift figure. The other 15 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of -8 percentage points is the difference between those two pass rates over the 10 comparable cases. 8 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.