Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Long-running iterative development loops with pacing control and verifiable progress. Use when tasks require multiple iterations, many discrete steps, or periodic reflection with clear checkpoints; avoid for simple one-shot tasks or quick fixes.
.claude/skills/dicklesworthstone-ralph-wiggum/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-07 | ✗→✓ | ▲ Improved | -68% | 0% |
| case-08 | ✗→✓ | ▲ Improved | -40% | 0% |
| case-09 | ✗→✓ | ▲ Improved | -42% | 0% |
| case-10 | ✗→✓ | ▲ Improved | -70% | 0% |
| case-11 | ✗→✓ | ▲ Improved | -50% | 0% |
Use the ralph_start tool to begin a loop:
ralph_start({
name: "loop-name",
taskContent: "# Task\n\n## Goals\n- Goal 1\n\n## Checklist\n- [ ] Item 1\n- [ ] Item 2",
maxIterations: 50, // Default: 50
itemsPerIteration: 3, // Optional: suggest N items per turn
reflectEvery: 10 // Optional: reflect every N iterations
}).ralph/<name>.md with the task content. The tool does NOT create this file—you must write it yourself using the Write tool.ralph_done to proceed to the next iteration.<promise>COMPLETE</promise> when finished./ralph start <name|path> - Start a new loop./ralph resume <name> - Resume loop./ralph stop - Pause loop (when agent idle)./ralph-stop - Stop active loop (idle only)./ralph status - Show loops./ralph list --archived - Show archived loops./ralph archive <name> - Move loop to archive./ralph clean [--all] - Clean completed loops./ralph cancel <name> - Delete loop./ralph nuke [--yes] - Delete all .ralph data.Press ESC to interrupt streaming, send a normal message to resume, and run /ralph-stop when idle to end the loop.
markdown# Task Title Brief description. ## Goals - Goal 1 - Goal 2 ## Checklist - [ ] Item 1 - [ ] Item 2 - [x] Completed item ## Verification - Evidence, commands run, or file paths ## Notes (Update with progress, decisions, blockers)
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-06 | pass→pass | 17,894 | 13,590 | -24% | 1 | 1 | 0% | 3,785 | 3,487 | -8% | 0 | 0 | — |
case-01 | fail→fail | 8,248 | 20,214 | +145% | 1 | 1 | 0% | 1,380 | 1,825 | +32% | 0 | 0 | — |
case-02 | fail→fail | 6,939 | 6,797 | -2% | 1 | 1 | 0% | 651 | 1,121 | +72% | 0 | 0 | — |
case-03 | fail→fail | 4,039 | 4,430 | +10% | 1 | 1 | 0% | 253 | 787 | +211% | 0 | 0 | — |
case-04 | pass→fail | 5,934 | 3,827 | -36% | 1 | 1 | 0% | 1,077 | 782 | -27% | 0 | 0 | — |
case-05 | pass→pass | 9,755 | 22,118 | +127% | 1 | 1 | 0% | 1,403 | 4,120 | +194% | 0 | 0 | — |
case-07 | fail→pass | 15,034 | 1,195 | -92% | 1 | 1 | 0% | 2,429 | 778 | -68% | 0 | 0 | — |
case-08 | fail→pass | 9,683 | 2,198 | -77% | 1 | 1 | 0% | 1,593 | 951 | -40% | 0 | 0 | — |
case-09 | fail→pass | 8,925 | 1,407 | -84% | 1 | 1 | 0% | 1,401 | 812 | -42% | 0 | 0 | — |
case-10 | fail→pass | 18,577 | 1,877 | -90% | 1 | 1 | 0% | 3,036 | 919 | -70% | 0 | 0 | — |
case-11 | fail→pass | 12,634 | 2,446 | -81% | 1 | 1 | 0% | 1,867 | 937 | -50% | 0 | 0 | — |
case-12 | fail→pass | 6,694 | 3,020 | -55% | 1 | 1 | 0% | 1,146 | 1,064 | -7% | 0 | 0 | — |
case-13 | fail→pass | 9,327 | 4,373 | -53% | 1 | 1 | 0% | 1,537 | 1,355 | -12% | 0 | 0 | — |
case-14 | fail→pass | 8,387 | 2,544 | -70% | 1 | 1 | 0% | 1,274 | 958 | -25% | 0 | 0 | — |
case-15 | pass→pass | 10,220 | 2,746 | -73% | 1 | 1 | 0% | 1,573 | 999 | -36% | 0 | 0 | — |
case-16 | fail→pass | 7,279 | 2,060 | -72% | 1 | 1 | 0% | 1,283 | 908 | -29% | 0 | 0 | — |
case-17 | fail→pass | 7,769 | 1,845 | -76% | 1 | 1 | 0% | 1,431 | 899 | -37% | 0 | 0 | — |
case-18 | fail→pass | 5,730 | 1,268 | -78% | 1 | 1 | 0% | 861 | 765 | -11% | 0 | 0 | — |
case-19 | pass→pass | 10,105 | 1,329 | -87% | 1 | 1 | 0% | 1,598 | 761 | -52% | 0 | 0 | — |
case-20 | fail→pass | 9,972 | 2,073 | -79% | 1 | 1 | 0% | 1,493 | 909 | -39% | 0 | 0 | — |
case-21 | fail→pass | 5,720 | 2,777 | -51% | 1 | 1 | 0% | 954 | 1,065 | +12% | 0 | 0 | — |
case-22 | pass→pass | 7,662 | 4,850 | -37% | 1 | 1 | 0% | 596 | 1,333 | +124% | 0 | 0 | — |
case-23 | fail→pass | 5,534 | 2,001 | -64% | 1 | 1 | 0% | 869 | 885 | +2% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 19 counted toward the lift figure. The other 4 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +57 percentage points is the difference between those two pass rates over the 19 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.