Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Run STORM phase 4 — article polishing. This skill should be used when the user asks to "polish the storm article", "finalize the article", or invokes /storm:polish. Adds a summary section, removes duplicate content, and verifies citation integrity.
.claude/skills/fradser-storm-polish/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | -2% | 0% |
| case-05 | ✗→✓ | ▲ Improved | -52% | 0% |
| case-09 | ✗→✓ | ▲ Improved | -36% | 0% |
| case-10 | ✗→✓ | ▲ Improved | -32% | 0% |
| case-16 | ✗→✓ | ▲ Improved | -9% | 0% |
Phase 4 of the STORM pipeline. Adds a summary/intro section, removes duplicate content across sections, and verifies every inline [n] citation resolves to a References entry and vice versa.
storm-engine via the Skill tool.article.md MUST exist (phase 3 complete). If absent, stop and instruct the user to run /storm:write first.This phase is complete iff article-polished.md exists. If --force is not set and it exists, skip and exit early.
article.md, outline.md, and research/sources.json.outline.md had an "Introduction" or "Summary" placeholder, write a summary section synthesizing the article's main points (1-2 paragraphs). Do not introduce new claims or citations not already in the body.--remove-duplicate is explicit but the behavior is the default) — detect near-duplicate paragraphs across sections and remove the later occurrence, keeping the one in the more topically-appropriate section.[n] keys present in the body. Append a ## References section listing exactly those sources, numbered to match, each as n. title — url (accessed YYYY-MM-DD). Drop any [n] in the body that has no source (replace with <!-- TODO: missing source -->). Drop any source not cited (do not list uncited sources in References).article-polished.md.run-config.json: phases.polish = "completed", final word count, source count, any integrity warnings.Report: final word count, number of cited sources, number of duplicate paragraphs removed, any integrity warnings (missing sources / TODO sections), and the absolute path to article-polished.md.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 3,617 | 7,267 | +101% | 1 | 1 | 0% | 288 | 616 | +114% | 0 | 0 | — |
case-02 | fail→fail | 38,556 | 4,005 | -90% | 1 | 1 | 0% | 6,415 | 671 | -90% | 0 | 0 | — |
case-03 | fail→fail | 11,680 | 6,683 | -43% | 1 | 1 | 0% | 1,893 | 888 | -53% | 0 | 0 | — |
case-04 | fail→pass | 6,171 | 3,419 | -45% | 1 | 1 | 0% | 978 | 957 | -2% | 0 | 0 | — |
case-05 | fail→pass | 11,381 | 2,478 | -78% | 1 | 1 | 0% | 1,743 | 840 | -52% | 0 | 0 | — |
case-06 | pass→pass | 6,903 | 5,055 | -27% | 1 | 1 | 0% | 1,119 | 1,242 | +11% | 0 | 0 | — |
case-07 | fail→fail | 12,505 | 8,602 | -31% | 1 | 1 | 0% | 1,929 | 1,835 | -5% | 0 | 0 | — |
case-08 | pass→pass | 2,269 | 2,489 | +10% | 1 | 1 | 0% | 286 | 856 | +199% | 0 | 0 | — |
case-09 | fail→pass | 9,290 | 2,662 | -71% | 1 | 1 | 0% | 1,300 | 832 | -36% | 0 | 0 | — |
case-10 | fail→pass | 10,123 | 3,840 | -62% | 1 | 1 | 0% | 1,549 | 1,060 | -32% | 0 | 0 | — |
case-11 | pass→pass | 8,926 | 4,381 | -51% | 1 | 1 | 0% | 1,425 | 1,221 | -14% | 0 | 0 | — |
case-16 | fail→pass | 6,049 | 2,497 | -59% | 1 | 1 | 0% | 930 | 845 | -9% | 0 | 0 | — |
case-12 | fail→pass | 11,233 | 3,472 | -69% | 1 | 1 | 0% | 2,039 | 1,106 | -46% | 0 | 0 | — |
case-13 | fail→pass | 8,006 | 2,709 | -66% | 1 | 1 | 0% | 1,432 | 869 | -39% | 0 | 0 | — |
case-14 | fail→pass | 14,773 | 5,252 | -64% | 1 | 1 | 0% | 2,658 | 1,328 | -50% | 0 | 0 | — |
case-15 | fail→fail | 10,641 | 2,651 | -75% | 1 | 1 | 0% | 1,627 | 812 | -50% | 0 | 0 | — |
case-17 | fail→pass | 17,200 | 3,974 | -77% | 1 | 1 | 0% | 3,312 | 1,146 | -65% | 0 | 0 | — |
case-18 | fail→fail | 12,418 | 1,633 | -87% | 1 | 1 | 0% | 1,793 | 611 | -66% | 0 | 0 | — |
case-19 | pass→pass | 10,485 | 3,267 | -69% | 1 | 1 | 0% | 1,596 | 1,008 | -37% | 0 | 0 | — |
case-20 | pass→pass | 7,911 | 2,156 | -73% | 1 | 1 | 0% | 1,276 | 779 | -39% | 0 | 0 | — |
case-21 | fail→fail | 8,488 | 2,028 | -76% | 1 | 1 | 0% | 1,317 | 762 | -42% | 0 | 0 | — |
case-22 | fail→fail | 2,636 | 14,500 | +450% | 1 | 1 | 0% | 410 | 2,910 | +610% | 0 | 0 | — |
case-23 | fail→fail | 5,922 | 9,514 | +61% | 1 | 1 | 0% | 899 | 1,916 | +113% | 0 | 0 | — |
case-24 | pass→pass | 6,169 | 6,419 | +4% | 1 | 1 | 0% | 1,261 | 1,739 | +38% | 0 | 0 | — |
case-25 | pass→pass | 8,348 | 10,237 | +23% | 1 | 1 | 0% | 1,535 | 2,428 | +58% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 25 cases were attempted, and 22 counted toward the lift figure. The other 3 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +36 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.