Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Start a ralph style looping session
.claude/skills/sterlingcrispin-initialize-project-documentation-from-system-plan/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-07 | ✗→✓ | ▲ Improved | 5% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 32% | 0% |
| case-16 | ✗→✓ | ▲ Improved | 82% | 0% |
| case-04 | ✓→✗ | ▼ Worse | 18% | 0% |
| case-17 | ✓→✗ | ▼ Worse | -26% | 0% |
You are initializing a new project for agent-driven development. The user has provided a current_SYSTEM-PLAN.md file in the current working directory that describes the system architecture and goals.
Read current_SYSTEM-PLAN.md and create two companion documents:
These three current_* files will serve as the foundation for an automated agent that works in a continuous development loop, picking up tasks and making progress incrementally.
First, read current_SYSTEM-PLAN.md in the current directory to understand:
Write a comprehensive PRD with testable requirements extracted from the system plan. Structure it as:
markdown# [Project Name] - Product Requirements Document **Last Updated:** [today's date] --- ## Requirements
{ "category": "functional|integration|architecture|performance|usability|maintainability|milestone|constraint", "description": "Clear, testable requirement statement", "steps": "Step 1 to verify this requirement", "Step 2...", "Step 3...", "Step 4...", "Step 5..." ], "passes": false }, ... ]
## Summary Statistics
| Category | Count |
|----------|-------|
| Functional | X |
| Integration | X |
| ... | ... |
| **Total** | **X** |
## Development Stages Mapping
Map requirements to development stages from the system plan.
## Open Questions
List any ambiguities or decisions that need user input."passes": falsefunctional, integration, architecture, performance, usability, maintainability, milestone, constraintWrite a progress journal structured for continuous agent development:
markdown# [Project Name] Progress Journal **Purpose:** Track progress on implementation. Useful for context continuity between sessions. **Last Updated:** [today's date] --- ## Current Status: [STAGE NAME] NOT STARTED Brief description of where we are. --- ## Project Context Quick Reference ### What We Have - Key existing assets (data, models, infrastructure) ### Key Files - List important files and their purposes ### Core Design Decisions (from plan) 1. Decision 1 2. Decision 2 ... --- ## Development Stages Checklist ### Stage 1: [Name] **Goal:** [from system plan] - [ ] Task 1 - [ ] Task 2 - [ ] Task 3 **Validation:** [how to know this stage is complete] --- ### Stage 2: [Name] ... --- ## Session Log ### [Date] - Session 1 **Work Done:** - (empty - will be filled as work progresses) **Next Session Should:** - Start with first task from Stage 1 --- ## Architecture Quick Reference [Include a simplified architecture diagram or component list from the system plan] --- ## Key Interfaces (for quick reference) [Extract key interfaces/APIs from the system plan that the agent will need to implement] --- ## Notes & Decisions Log *Add notes here as implementation decisions are made* --- ## Open Questions [Copy from system plan or identify new ones]
After creating both files:
current_PRD.md and current_PROGRESS.md| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 1,977 | 4,767 | +141% | 1 | 1 | 0% | 295 | 1,469 | +398% | 0 | 0 | — |
case-02 | fail→fail | 8,854 | 3,792 | -57% | 1 | 1 | 0% | 1,344 | 1,348 | +0% | 0 | 0 | — |
case-03 | fail→fail | 26,422 | 4,663 | -82% | 1 | 1 | 0% | 4,867 | 1,381 | -72% | 0 | 0 | — |
case-04 | pass→fail | 7,451 | 5,910 | -21% | 1 | 1 | 0% | 1,401 | 1,648 | +18% | 0 | 0 | — |
case-05 | fail→fail | 4,149 | 1,736 | -58% | 1 | 1 | 0% | 648 | 1,458 | +125% | 0 | 0 | — |
case-06 | pass→pass | 6,168 | 2,717 | -56% | 1 | 1 | 0% | 1,061 | 1,635 | +54% | 0 | 0 | — |
case-07 | fail→pass | 8,731 | 1,973 | -77% | 1 | 1 | 0% | 1,437 | 1,514 | +5% | 0 | 0 | — |
case-08 | fail→pass | 7,937 | 3,177 | -60% | 1 | 1 | 0% | 1,313 | 1,734 | +32% | 0 | 0 | — |
case-09 | fail→fail | 6,828 | 3,310 | -52% | 1 | 1 | 0% | 1,077 | 1,729 | +61% | 0 | 0 | — |
case-10 | fail→fail | 11,603 | 7,383 | -36% | 1 | 1 | 0% | 1,867 | 2,443 | +31% | 0 | 0 | — |
case-11 | fail→fail | 6,759 | 1,628 | -76% | 1 | 1 | 0% | 1,131 | 1,433 | +27% | 0 | 0 | — |
case-12 | fail→fail | 7,224 | 1,972 | -73% | 1 | 1 | 0% | 1,212 | 1,488 | +23% | 0 | 0 | — |
case-13 | pass→pass | 5,405 | 2,183 | -60% | 1 | 1 | 0% | 903 | 1,545 | +71% | 0 | 0 | — |
case-14 | pass→pass | 6,209 | 2,932 | -53% | 1 | 1 | 0% | 1,057 | 1,627 | +54% | 0 | 0 | — |
case-15 | pass→pass | 8,805 | 2,604 | -70% | 1 | 1 | 0% | 1,560 | 1,575 | +1% | 0 | 0 | — |
case-16 | fail→pass | 6,115 | 1,977 | -68% | 1 | 1 | 0% | 844 | 1,536 | +82% | 0 | 0 | — |
case-17 | pass→fail | 11,714 | 1,553 | -87% | 1 | 1 | 0% | 1,896 | 1,398 | -26% | 0 | 0 | — |
case-18 | pass→pass | 8,533 | 3,274 | -62% | 1 | 1 | 0% | 1,314 | 1,715 | +31% | 0 | 0 | — |
case-19 | fail→fail | 4,959 | 2,394 | -52% | 1 | 1 | 0% | 658 | 1,545 | +135% | 0 | 0 | — |
case-20 | pass→fail | 9,496 | 3,688 | -61% | 1 | 1 | 0% | 1,214 | 1,371 | +13% | 0 | 0 | — |
case-21 | pass→fail | 27,702 | 32,001 | +16% | 1 | 1 | 0% | 5,397 | 7,381 | +37% | 0 | 0 | — |
case-22 | pass→fail | 10,331 | 9,431 | -9% | 1 | 1 | 0% | 2,110 | 1,476 | -30% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 16 counted toward the lift figure. The other 6 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of -9 percentage points is the difference between those two pass rates over the 16 comparable cases. 5 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.