Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Build per-chapter (H2) writing briefs (NO PROSE) so the final survey reads like a paper (chapter leads + cross-H3 coherence) without inflating the ToC. **Trigger**: chapter briefs, H2 briefs, chapter lead plan, section intent, 章节意图, 章节导读, H2 卡片. **Use when**: `outline/outline.yml` + `outline/subsection_briefs.jsonl` exist and you want thicker chapters (fewer headings, more logic).
.claude/skills/willoscar-chapter-briefs/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-07 | ✗→✓ | ▲ Improved | 79% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 33% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 98% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 43% | 0% |
| case-11 | ✗→✓ | ▲ Improved | -10% | 0% |
outline/outline.yml + outline/subsection_briefs.jsonl exist and you want thicker chapters (fewer headings, more logic).Purpose: turn each H2 chapter that contains H3 subsections into a chapter-level writing card so the writer can:
This artifact is internal intent, not reader-facing prose.
Why this matters for writing quality:
throughline and lead_paragraph_plan as decision constraints, not copyable sentences.outline/outline.ymloutline/subsection_briefs.jsonlGOAL.mdoutline/chapter_briefs.jsonloutline/chapter_briefs.jsonl)JSONL (one object per H2 chapter that has H3 subsections).
Required fields:
section_id, section_titlesubsections (list of {sub_id,title} in outline order)synthesis_mode (one of: clusters, timeline, tradeoff_matrix, case_study, tension_resolution)synthesis_preview (1–2 bullets; how the chapter will synthesize across H3 without template-y “Taken together…”)throughline (3–6 bullets)key_contrasts (2–6 bullets; pull from each H3 contrast_hook when available)lead_paragraph_plan (2–3 bullets; plan only, not prose)bridge_terms (5–12 tokens; union of H3 bridge terms)The writer uses outline/chapter_briefs.jsonl to draft sections/S<sec_id>_lead.md (body-only; no headings).
Contract (paper-like, no new facts):
key_contrasts / bridge_terms as handles (not templates) so the chapter reads coherent without repeating "Taken together" everywhere.GOAL.md to pin scope/audience, and inject that constraint into the chapter throughline.outline/outline.yml and list H2 chapters that have H3 subsections.outline/subsection_briefs.jsonl and group briefs by section_id.outline/chapter_briefs.jsonl.TODO/…/(placeholder)/template instructions).throughline and key_contrasts are chapter-specific (not copy/paste generic).lead_paragraph_plan bullets explicitly preview 2–3 comparison axes and how the H3 subsections partition them (no generic chapter-intro boilerplate).uv run python .codex/skills/chapter-briefs/scripts/run.py --helpuv run python .codex/skills/chapter-briefs/scripts/run.py --workspace <workspace>--workspace <dir>--unit-id <U###>--inputs <semicolon-separated>--outputs <semicolon-separated>--checkpoint <C#>uv run python .codex/skills/chapter-briefs/scripts/run.py --workspace <workspace>uv run python .codex/skills/chapter-briefs/scripts/run.py --workspace <workspace> --inputs "outline/outline.yml;outline/subsection_briefs.jsonl;GOAL.md" --outputs "outline/chapter_briefs.jsonl"When you are satisfied with chapter briefs, create:
outline/chapter_briefs.refined.okThis is an explicit "I reviewed/refined this" signal:
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 15,923 | 5,932 | -63% | 1 | 1 | 0% | 2,530 | 1,426 | -44% | 0 | 0 | — |
case-02 | fail→fail | 34,063 | 4,531 | -87% | 1 | 1 | 0% | 5,558 | 1,462 | -74% | 0 | 0 | — |
case-03 | fail→fail | 5,170 | 5,681 | +10% | 1 | 1 | 0% | 239 | 1,456 | +509% | 0 | 0 | — |
case-04 | pass→pass | 6,630 | 13,392 | +102% | 1 | 1 | 0% | 992 | 3,240 | +227% | 0 | 0 | — |
case-05 | pass→fail | 21,120 | 4,774 | -77% | 1 | 1 | 0% | 3,543 | 1,401 | -60% | 0 | 0 | — |
case-06 | fail→fail | 3,917 | 4,371 | +12% | 1 | 1 | 0% | 146 | 1,385 | +849% | 0 | 0 | — |
case-07 | fail→pass | 5,504 | 2,047 | -63% | 1 | 1 | 0% | 867 | 1,552 | +79% | 0 | 0 | — |
case-08 | fail→pass | 14,493 | 2,498 | -83% | 1 | 1 | 0% | 1,210 | 1,615 | +33% | 0 | 0 | — |
case-09 | fail→pass | 4,792 | 1,762 | -63% | 1 | 1 | 0% | 765 | 1,518 | +98% | 0 | 0 | — |
case-10 | fail→pass | 6,929 | 2,538 | -63% | 1 | 1 | 0% | 1,173 | 1,679 | +43% | 0 | 0 | — |
case-11 | fail→pass | 10,334 | 2,786 | -73% | 1 | 1 | 0% | 1,863 | 1,684 | -10% | 0 | 0 | — |
case-12 | fail→pass | 14,046 | 6,879 | -51% | 1 | 1 | 0% | 2,228 | 2,231 | +0% | 0 | 0 | — |
case-13 | fail→pass | 10,807 | 6,247 | -42% | 1 | 1 | 0% | 1,674 | 2,149 | +28% | 0 | 0 | — |
case-14 | pass→fail | 36,361 | 8,332 | -77% | 1 | 1 | 0% | 1,412 | 1,824 | +29% | 0 | 0 | — |
case-15 | fail→fail | 10,949 | 2,900 | -74% | 1 | 1 | 0% | 1,841 | 1,728 | -6% | 0 | 0 | — |
case-16 | fail→pass | 7,579 | 26,874 | +255% | 1 | 1 | 0% | 842 | 1,535 | +82% | 0 | 0 | — |
case-17 | fail→pass | 7,504 | 3,733 | -50% | 1 | 1 | 0% | 1,349 | 1,914 | +42% | 0 | 0 | — |
case-18 | pass→pass | 8,519 | 4,547 | -47% | 1 | 1 | 0% | 1,366 | 1,767 | +29% | 0 | 0 | — |
case-19 | fail→pass | 11,868 | 4,459 | -62% | 1 | 1 | 0% | 1,737 | 1,489 | -14% | 0 | 0 | — |
case-20 | fail→pass | 7,421 | 1,642 | -78% | 1 | 1 | 0% | 372 | 1,432 | +285% | 0 | 0 | — |
case-21 | pass→pass | 10,444 | 2,657 | -75% | 1 | 1 | 0% | 1,582 | 1,582 | 0% | 0 | 0 | — |
case-22 | fail→pass | 13,906 | 1,487 | -89% | 1 | 1 | 0% | 2,600 | 1,415 | -46% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 16 counted toward the lift figure. The other 6 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +45 percentage points is the difference between those two pass rates over the 16 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
The publisher has shipped newer versions since this run, so these numbers describe v1, not the version currently listed.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.