Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Create per-subsection evidence packs (NO PROSE): claim candidates, concrete comparisons, evaluation protocol, limitations, plus citation-backed evidence snippets with provenance. **Trigger**: evidence draft, evidence pack, claim candidates, concrete comparisons, evidence snippets, provenance, 证据草稿, 证据包, 可引用事实.
.claude/skills/willoscar-evidence-draft/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | 32% | 0% |
| case-13 | ✗→✓ | ▲ Improved | -10% | 0% |
| case-10 | ✗→✓ | ▲ Improved | -15% | 0% |
| case-11 | ✗→✓ | ▲ Improved | -14% | 0% |
| case-12 | ✗→✓ | ▲ Improved | -6% | 0% |
Build deterministic outline/evidence_drafts.jsonl packs from briefs + notes + optional evidence bindings.
Compatibility mode is active: this migration preserves the existing JSONL contract while moving evidence-quality policy, sparse-evidence routing, and evaluation-anchor rules into references/ and assets/.
Always read:
references/overview.mdreferences/evidence_quality_policy.mdRead by task:
references/block_vs_downgrade.md when deciding whether thin evidence should block drafting or only downgrade claim strengthreferences/evaluation_anchor_rules.md when evaluation tokens, protocol context, or numeric claims are weakreferences/examples_sparse_evidence.md for evidence-thin pack calibrationreferences/source_text_hygiene.md when paper self-narration or generic result wrappers are leaking into pack snippets / claim candidatesMachine-readable assets:
assets/evidence_pack_schema.jsonassets/evidence_policy.jsonassets/source_text_hygiene.jsonassets/limitation-signals.json — shared polarity rules that keepresolved failures and positive improvements out of limitation slots
Required:
outline/subsection_briefs.jsonlpapers/paper_notes.jsonlcitations/ref.bibOptional but recommended:
papers/evidence_bank.jsonloutline/evidence_bindings.jsonlKeep the current output contract:
outline/evidence_drafts.jsonloutline/evidence_drafts/Use scripts/run.py only for:
blocking_missing / downgrade_signals / verify_fields materializationDo not treat run.py as the place for:
references/ / assets/Keep these stable:
claim_candidates must remain snippet-derivedconcrete_comparisons must remain genuinely two-sided; if one cluster has no usable highlight, drop the card and surface thin evidence upstream instead of fabricating an A-vs-B contrastcitations/ref.bibCurrent mode is reference-first with deterministic compatibility:
assets/evidence_policy.json defines pack thresholds and sparse-evidence routingassets/evidence_pack_schema.json documents/validates the stable pack shapeassets/source_text_hygiene.json owns this Skill's wrapper cleanup, while therepo-wide assets/limitation-signals.json owns limitation polarity across paper-notes, evidence-draft, and writer-context-pack
scripts/run.py still materializes the existing JSONL + Markdown outputs, but no longer pads sparse sections with generic caution proseuv run python .codex/skills/evidence-draft/scripts/run.py --workspace <workspace>When running in compatibility mode, scripts/run.py currently reads:
outline/subsection_briefs.jsonlpapers/paper_notes.jsonlcitations/ref.bibpapers/evidence_bank.jsonl and outline/evidence_bindings.jsonlassets/evidence_policy.json and assets/evidence_pack_schema.jsonuv run python .codex/skills/evidence-draft/scripts/run.py --workspace <workspace>--workspace <dir>--unit-id <id>--inputs <path1;path2>--outputs <path1;path2>--checkpoint <C*>uv run python .codex/skills/evidence-draft/scripts/run.py --workspace <workspace>assets/evidence_policy.json and references/block_vs_downgrade.md before changing Python.references/evaluation_anchor_rules.md and the policy asset.downgrade_signals and verify_fields rather than adding narrative caveats.| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 4,409 | 5,317 | +21% | 1 | 1 | 0% | 181 | 1,433 | +692% | 0 | 0 | — |
case-02 | fail→fail | 6,141 | 5,685 | -7% | 1 | 1 | 0% | 237 | 1,425 | +501% | 0 | 0 | — |
case-03 | fail→fail | 3,973 | 5,098 | +28% | 1 | 1 | 0% | 243 | 1,337 | +450% | 0 | 0 | — |
case-04 | fail→pass | 10,194 | 6,078 | -40% | 1 | 1 | 0% | 1,549 | 2,050 | +32% | 0 | 0 | — |
case-13 | fail→pass | 11,244 | 2,412 | -79% | 1 | 1 | 0% | 1,639 | 1,472 | -10% | 0 | 0 | — |
case-05 | pass→pass | 12,392 | 5,047 | -59% | 1 | 1 | 0% | 1,468 | 1,862 | +27% | 0 | 0 | — |
case-06 | pass→pass | 9,994 | 3,904 | -61% | 1 | 1 | 0% | 1,699 | 1,835 | +8% | 0 | 0 | — |
case-07 | pass→pass | 15,210 | 6,854 | -55% | 1 | 1 | 0% | 2,127 | 2,230 | +5% | 0 | 0 | — |
case-08 | pass→pass | 11,947 | 5,280 | -56% | 1 | 1 | 0% | 1,690 | 1,928 | +14% | 0 | 0 | — |
case-09 | pass→pass | 10,257 | 3,696 | -64% | 1 | 1 | 0% | 1,746 | 1,857 | +6% | 0 | 0 | — |
case-10 | fail→pass | 12,042 | 3,731 | -69% | 1 | 1 | 0% | 1,900 | 1,607 | -15% | 0 | 0 | — |
case-11 | fail→pass | 9,301 | 1,627 | -83% | 1 | 1 | 0% | 1,552 | 1,332 | -14% | 0 | 0 | — |
case-12 | fail→pass | 10,915 | 2,255 | -79% | 1 | 1 | 0% | 1,603 | 1,514 | -6% | 0 | 0 | — |
case-14 | fail→pass | 8,495 | 2,009 | -76% | 1 | 1 | 0% | 1,297 | 1,493 | +15% | 0 | 0 | — |
case-15 | fail→pass | 6,797 | 1,504 | -78% | 1 | 1 | 0% | 976 | 1,327 | +36% | 0 | 0 | — |
case-16 | fail→pass | 9,132 | 1,485 | -84% | 1 | 1 | 0% | 1,302 | 1,336 | +3% | 0 | 0 | — |
case-17 | fail→pass | 15,275 | 3,131 | -80% | 1 | 1 | 0% | 2,258 | 1,660 | -26% | 0 | 0 | — |
case-18 | fail→pass | 11,168 | 1,701 | -85% | 1 | 1 | 0% | 1,644 | 1,367 | -17% | 0 | 0 | — |
case-19 | fail→pass | 13,350 | 4,971 | -63% | 1 | 1 | 0% | 1,974 | 1,931 | -2% | 0 | 0 | — |
case-20 | pass→pass | 19,482 | 26,629 | +37% | 1 | 1 | 0% | 2,644 | 5,300 | +100% | 0 | 0 | — |
case-21 | fail→fail | 2,765 | 7,594 | +175% | 1 | 1 | 0% | 366 | 2,374 | +549% | 0 | 0 | — |
case-22 | fail→fail | 4,247 | 2,355 | -45% | 1 | 1 | 0% | 575 | 1,489 | +159% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 19 counted toward the lift figure. The other 3 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +50 percentage points is the difference between those two pass rates over the 19 comparable cases.
The publisher has shipped newer versions since this run, so these numbers describe v1, not the version currently listed.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.