Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Run STORM phase 2 — outline generation. This skill should be used when the user asks to "generate an outline for a storm article", "draft a wikipedia-style outline", or invokes /storm:outline. Produces a draft outline from parametric knowledge then refines it using the research information table.
.claude/skills/fradser-storm-outline/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-14 | ✗→✓ | ▲ Improved | -45% | 0% |
| case-08 | ✗→✓ | ▲ Improved | -71% | 0% |
| case-09 | ✗→✓ | ▲ Improved | -44% | 0% |
| case-11 | ✗→✓ | ▲ Improved | -59% | 0% |
| case-13 | ✗→✓ | ▲ Improved | -49% | 0% |
Phase 2 of the STORM pipeline. Drafts an outline from the model's parametric knowledge, then refines it using the research conversations to reflect what was actually learned.
storm-engine via the Skill tool.research/sources.json MUST exist (phase 1 complete). If absent, stop and instruct the user to run /storm:research first. Do not proceed with a parametric-only outline as the final artifact.This phase is complete iff outline.md exists and has ≥2 sections. If --force is not set and it exists, skip and exit early.
research/conversations.jsonl and research/sources.json. Concatenate the conversation histories as the refinement input.outline-draft.md from parametric knowledge alone. This is the model's prior structure for the topic. Use markdown ## Section headings.outline.md.polish phase, not here.outline.md has ≥2 sections. Update run-config.json: phases.outline = "completed", section count.Report: number of draft sections, number of refined sections, sections added/removed during refinement, and the path to outline.md.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-14 | fail→pass | 9,219 | 2,365 | -74% | 1 | 1 | 0% | 1,477 | 809 | -45% | 0 | 0 | — |
case-01 | fail→fail | 14,669 | 8,231 | -44% | 1 | 1 | 0% | 2,466 | 659 | -73% | 0 | 0 | — |
case-02 | fail→fail | 19,085 | 3,728 | -80% | 1 | 1 | 0% | 2,575 | 587 | -77% | 0 | 0 | — |
case-03 | fail→fail | 13,665 | 4,995 | -63% | 1 | 1 | 0% | 2,307 | 777 | -66% | 0 | 0 | — |
case-15 | pass→pass | 11,912 | 8,478 | -29% | 1 | 1 | 0% | 1,932 | 1,721 | -11% | 0 | 0 | — |
case-04 | fail→fail | 12,361 | 8,731 | -29% | 1 | 1 | 0% | 1,771 | 1,069 | -40% | 0 | 0 | — |
case-05 | pass→fail | 6,250 | 4,999 | -20% | 1 | 1 | 0% | 899 | 599 | -33% | 0 | 0 | — |
case-06 | fail→fail | 9,526 | 4,926 | -48% | 1 | 1 | 0% | 320 | 702 | +119% | 0 | 0 | — |
case-07 | fail→fail | 10,185 | 2,488 | -76% | 1 | 1 | 0% | 1,631 | 776 | -52% | 0 | 0 | — |
case-08 | fail→pass | 17,749 | 2,790 | -84% | 1 | 1 | 0% | 2,626 | 768 | -71% | 0 | 0 | — |
case-09 | fail→pass | 9,925 | 3,784 | -62% | 1 | 1 | 0% | 1,536 | 861 | -44% | 0 | 0 | — |
case-10 | pass→pass | 7,816 | 2,619 | -66% | 1 | 1 | 0% | 1,215 | 814 | -33% | 0 | 0 | — |
case-11 | fail→pass | 10,107 | 2,579 | -74% | 1 | 1 | 0% | 1,719 | 704 | -59% | 0 | 0 | — |
case-12 | pass→fail | 9,403 | 2,402 | -74% | 1 | 1 | 0% | 1,539 | 771 | -50% | 0 | 0 | — |
case-13 | fail→pass | 6,796 | 1,501 | -78% | 1 | 1 | 0% | 1,182 | 597 | -49% | 0 | 0 | — |
case-16 | pass→pass | 9,242 | 5,123 | -45% | 1 | 1 | 0% | 1,408 | 1,122 | -20% | 0 | 0 | — |
case-17 | fail→fail | 10,319 | 1,677 | -84% | 1 | 1 | 0% | 1,615 | 591 | -63% | 0 | 0 | — |
case-18 | fail→pass | 8,494 | 3,395 | -60% | 1 | 1 | 0% | 1,242 | 942 | -24% | 0 | 0 | — |
case-19 | fail→fail | 4,875 | 5,484 | +12% | 1 | 1 | 0% | 242 | 813 | +236% | 0 | 0 | — |
case-20 | fail→fail | 6,313 | 5,501 | -13% | 1 | 1 | 0% | 830 | 682 | -18% | 0 | 0 | — |
case-21 | fail→fail | 2,316 | 10,829 | +368% | 1 | 1 | 0% | 235 | 2,108 | +797% | 0 | 0 | — |
case-22 | pass→pass | 13,688 | 11,112 | -19% | 1 | 1 | 0% | 1,838 | 2,051 | +12% | 0 | 0 | — |
case-23 | pass→pass | 7,919 | 2,664 | -66% | 1 | 1 | 0% | 1,408 | 831 | -41% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 15 counted toward the lift figure. The other 8 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +17 percentage points is the difference between those two pass rates over the 15 comparable cases. 3 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.