Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Build per-subsection writing briefs (NO PROSE) so later drafting is driven by evidence and checkable comparison axes (not outline placeholders). **Trigger**: subsection briefs, writing cards, intent cards, H3 briefs, scope_rule, axes, clusters, 写作意图卡, 小节卡片, 段落计划.
.claude/skills/willoscar-subsection-briefs/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-05 | ✗→✓ | ▲ Improved | 13% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 15% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 89% | 0% |
| case-12 | ✗→✓ | ▲ Improved | -32% | 0% |
| case-15 | ✗→✓ | ▲ Improved | 55% | 0% |
outline/subsection_briefs.refined.ok freezes reviewed briefs only while the marker is newer than the briefs, declared inputs, domain packs, and generator. Stale markers are removed before backed-up regeneration.
Build deterministic H3 brief cards from outline + mapping + paper notes.
Compatibility mode is active: this skill keeps the current outline/subsection_briefs.jsonl field contract and paragraph-plan shape while moving phrase/domain logic into references/ and assets/.
scripts/run.py as the deterministic materializer.transition-weaver, writer-context-pack, and subsection-writer.Always read:
references/overview.mdRead by task:
thesis feels repetitive or copyable, read references/thesis_patterns.md.tension_statement is too generic, read references/tension_patterns.md.references/axis_catalog_generic.md and references/axis_catalog_llm_agents.md.references/bridge_terms.md.references/examples_good.md.Machine-readable assets:
assets/phrase_packs/thesis_patterns.jsonassets/phrase_packs/bridge_contrast.jsonassets/domain_packs/generic.jsonassets/domain_packs/llm_agents.jsonassets/domain_packs/embodied_ai.jsonassets/domain_packs/rag_evaluation.jsonassets/domain_packs/text_to_image.jsonThe script loads these packs first; patch them before changing Python when the issue is phrasing, domain routing, axis inventory, cluster purity, or lexical bridge coverage.
outline/outline.ymloutline/mapping.tsvpapers/paper_notes.jsonlGOAL.mdoutline/claim_evidence_matrix.mdoutline/subsection_briefs.jsonlRequired record shape remains compatibility-preserving:
sub_id, title, section_id, section_titlerq, thesis, scope_rule, axes, bridge_terms, contrast_hook, tension_statementevaluation_anchor_minimal, required_evidence_fields, clustersparagraph_plan, evidence_level_summary, generated_atrun.py Should Doassets/.run.py Should Not Dothesis/tension_statement conservative and let downstream evidence skills strengthen the subsection.bridge_terms to surface concrete lexical handles that later evidence/ranking stages can still match (OOD, sim-to-real, world model, failure detector, specific benchmark families), not only generic axis names.cluster_rules over ad-hoc bootstrap overlaps when the mapped set is already large enough to support disjoint clusters.When running in compatibility mode, scripts/run.py currently reads:
outline/outline.yml for section/subsection structureoutline/mapping.tsv for paper-to-subsection coveragepapers/paper_notes.jsonl for structured evidenceGOAL.md for topic/domain cuesoutline/claim_evidence_matrix.md as optional supporting context when presentuv run python .codex/skills/subsection-briefs/scripts/run.py --workspace <workspace>--workspace <dir>--unit-id <id>--inputs <a;b;...>--outputs <a;b;...>--checkpoint <C*>uv run python .codex/skills/subsection-briefs/scripts/run.py --workspace <workspace>GOAL.md and the asset packs before changing the script.papers/paper_notes.jsonl is thin, reroute to note extraction rather than inventing axes.| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-13 | fail→fail | 5,767 | 4,330 | -25% | 1 | 1 | 0% | 221 | 1,349 | +510% | 0 | 0 | — |
case-02 | fail→fail | 18,959 | 5,605 | -70% | 1 | 1 | 0% | 3,125 | 1,525 | -51% | 0 | 0 | — |
case-01 | fail→fail | 21,685 | 6,616 | -69% | 1 | 1 | 0% | 3,653 | 1,449 | -60% | 0 | 0 | — |
case-03 | fail→fail | 8,831 | 4,844 | -45% | 1 | 1 | 0% | 1,435 | 1,369 | -5% | 0 | 0 | — |
case-04 | pass→pass | 11,414 | 2,863 | -75% | 1 | 1 | 0% | 1,612 | 1,648 | +2% | 0 | 0 | — |
case-05 | fail→pass | 13,243 | 6,512 | -51% | 1 | 1 | 0% | 1,816 | 2,060 | +13% | 0 | 0 | — |
case-06 | pass→pass | 14,163 | 2,859 | -80% | 1 | 1 | 0% | 2,114 | 1,638 | -23% | 0 | 0 | — |
case-07 | fail→fail | 11,860 | 3,703 | -69% | 1 | 1 | 0% | 1,767 | 1,828 | +3% | 0 | 0 | — |
case-08 | fail→pass | 9,903 | 3,516 | -64% | 1 | 1 | 0% | 1,521 | 1,750 | +15% | 0 | 0 | — |
case-09 | fail→pass | 5,402 | 1,995 | -63% | 1 | 1 | 0% | 808 | 1,529 | +89% | 0 | 0 | — |
case-10 | pass→pass | 13,997 | 2,531 | -82% | 1 | 1 | 0% | 1,994 | 1,581 | -21% | 0 | 0 | — |
case-11 | pass→pass | 14,033 | 3,639 | -74% | 1 | 1 | 0% | 1,959 | 1,768 | -10% | 0 | 0 | — |
case-12 | fail→pass | 22,580 | 6,900 | -69% | 1 | 1 | 0% | 3,440 | 2,330 | -32% | 0 | 0 | — |
case-14 | fail→fail | 9,348 | 5,138 | -45% | 1 | 1 | 0% | 710 | 1,396 | +97% | 0 | 0 | — |
case-15 | fail→pass | 17,536 | 1,443 | -92% | 1 | 1 | 0% | 885 | 1,375 | +55% | 0 | 0 | — |
case-16 | pass→pass | 6,425 | 2,361 | -63% | 1 | 1 | 0% | 950 | 1,545 | +63% | 0 | 0 | — |
case-17 | fail→pass | 12,291 | 1,779 | -86% | 1 | 1 | 0% | 1,888 | 1,449 | -23% | 0 | 0 | — |
case-18 | pass→pass | 8,589 | 1,876 | -78% | 1 | 1 | 0% | 1,247 | 1,474 | +18% | 0 | 0 | — |
case-19 | fail→pass | 11,008 | 2,426 | -78% | 1 | 1 | 0% | 1,513 | 1,548 | +2% | 0 | 0 | — |
case-20 | fail→pass | 7,951 | 2,066 | -74% | 1 | 1 | 0% | 1,175 | 1,540 | +31% | 0 | 0 | — |
case-21 | pass→pass | 6,944 | 1,504 | -78% | 1 | 1 | 0% | 923 | 1,413 | +53% | 0 | 0 | — |
case-22 | fail→pass | 9,042 | 1,823 | -80% | 1 | 1 | 0% | 1,160 | 1,448 | +25% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 16 counted toward the lift figure. The other 6 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +41 percentage points is the difference between those two pass rates over the 16 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.