Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Rewrite limitation passages so the paper keeps limitations without falling into count-based slot phrases (e.g., \"Two limitations…\") across many H3s. **Trigger**: limitation weaver, rewrite limitations, remove two limitations, 去Two limitations, 局限改写, caveat rewrite.
.claude/skills/willoscar-limitation-weaver/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-10 | ✗→✓ | ▲ Improved | -10% | 0% |
| case-11 | ✗→✓ | ▲ Improved | -51% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 25% | 0% |
| case-16 | ✗→✓ | ▲ Improved | 9% | 0% |
| case-18 | ✗→✓ | ▲ Improved | -5% | 0% |
Purpose: keep survey-grade intellectual honesty without triggering a strong generator-voice tell:
This is not about removing limitations. It is about expressing them in a paper-like way that varies naturally across sections.
Required:
output/WRITER_SELFLOOP_TODO.md (Style Smells section)sections/S<sub_id>.md filesOptional (helps keep limitations grounded):
outline/writer_context_packs.jsonl (use failures_limitations / limitation_hooks / verify_fields when present)output/WRITER_SELFLOOP_TODO.md (Style Smells) to locate the exact sections/S*.md files to rewrite.outline/writer_context_packs.jsonl to keep limitations grounded in the subsection's evidence boundary (no guessing).sections/S<sub_id>.md files (still body-only; no headings)textYou are editing the limitation content of a survey subsection. Goal: - preserve the subsection-specific limitation(s) - remove count-based opener slots and repetitive cadence - keep limitations tied to the protocol/evidence boundary (what changes interpretation) Constraints: - do not invent facts - do not add/remove/move citation keys - do not weaken the section by deleting real limitations
Two limitations stand out. First, ... Second, ...Three key takeaways are ...Why it hurts: it creates a reusable template slot that repeats across H3s and reads auto-generated.
1) Fold caveat into a contrast paragraph (preferred)
2) Single caveat paragraph without counting
3) Verification-target framing (when evidence is abstract-only / underspecified)
Bad:
Two limitations temper strong conclusions. First, budgets differ. Second, ablations are missing.Better (folded into contrast):
...; however, reported budgets and retry policies vary widely, which makes head-to-head comparisons fragile unless those constraints are normalized.Better (single caveat paragraph):
These results hinge on under-specified verification and retry policies; this matters because success rates can shift substantially along the success–cost frontier.writer-selfloop remains PASS and Style Smells shrink.| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 4,707 | 5,068 | +8% | 1 | 1 | 0% | 265 | 1,098 | +314% | 0 | 0 | — |
case-02 | fail→fail | 5,050 | 4,361 | -14% | 1 | 1 | 0% | 327 | 1,033 | +216% | 0 | 0 | — |
case-03 | fail→fail | 3,163 | 8,695 | +175% | 1 | 1 | 0% | 544 | 2,274 | +318% | 0 | 0 | — |
case-04 | fail→fail | 5,473 | 9,619 | +76% | 1 | 1 | 0% | 857 | 1,963 | +129% | 0 | 0 | — |
case-05 | fail→fail | 5,529 | 6,275 | +13% | 1 | 1 | 0% | 763 | 1,793 | +135% | 0 | 0 | — |
case-15 | pass→pass | 11,206 | 6,248 | -44% | 1 | 1 | 0% | 1,656 | 1,699 | +3% | 0 | 0 | — |
case-06 | pass→pass | 7,362 | 7,692 | +4% | 1 | 1 | 0% | 1,111 | 2,028 | +83% | 0 | 0 | — |
case-07 | pass→pass | 10,193 | 5,253 | -48% | 1 | 1 | 0% | 1,441 | 1,515 | +5% | 0 | 0 | — |
case-08 | pass→pass | 7,898 | 4,338 | -45% | 1 | 1 | 0% | 1,116 | 1,460 | +31% | 0 | 0 | — |
case-09 | fail→fail | 8,647 | 3,509 | -59% | 1 | 1 | 0% | 1,391 | 1,341 | -4% | 0 | 0 | — |
case-10 | fail→pass | 8,552 | 2,231 | -74% | 1 | 1 | 0% | 1,281 | 1,154 | -10% | 0 | 0 | — |
case-11 | fail→pass | 13,463 | 1,915 | -86% | 1 | 1 | 0% | 2,305 | 1,138 | -51% | 0 | 0 | — |
case-12 | fail→pass | 8,477 | 5,058 | -40% | 1 | 1 | 0% | 1,239 | 1,552 | +25% | 0 | 0 | — |
case-13 | pass→pass | 8,780 | 5,477 | -38% | 1 | 1 | 0% | 1,296 | 1,603 | +24% | 0 | 0 | — |
case-14 | pass→pass | 7,516 | 5,848 | -22% | 1 | 1 | 0% | 1,104 | 1,642 | +49% | 0 | 0 | — |
case-16 | fail→pass | 11,137 | 7,185 | -35% | 1 | 1 | 0% | 1,634 | 1,778 | +9% | 0 | 0 | — |
case-17 | pass→pass | 2,806 | 3,759 | +34% | 1 | 1 | 0% | 427 | 1,419 | +232% | 0 | 0 | — |
case-18 | fail→pass | 10,521 | 5,635 | -46% | 1 | 1 | 0% | 1,750 | 1,658 | -5% | 0 | 0 | — |
case-19 | pass→pass | 9,453 | 3,957 | -58% | 1 | 1 | 0% | 1,437 | 1,359 | -5% | 0 | 0 | — |
case-20 | pass→pass | 7,040 | 7,815 | +11% | 1 | 1 | 0% | 999 | 1,782 | +78% | 0 | 0 | — |
case-21 | fail→pass | 12,437 | 2,471 | -80% | 1 | 1 | 0% | 2,035 | 1,162 | -43% | 0 | 0 | — |
case-22 | fail→pass | 6,478 | 2,251 | -65% | 1 | 1 | 0% | 924 | 1,154 | +25% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 20 counted toward the lift figure. The other 2 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +32 percentage points is the difference between those two pass rates over the 20 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.