Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Build a section-by-section claim–evidence matrix (`outline/claim_evidence_matrix.md`) from the outline and paper notes. **Trigger**: claim–evidence matrix, evidence mapping, 证据矩阵, 主张-证据对齐. **Use when**: 写 prose 之前需要把每个小节的可检验主张与证据来源显式化(outline + paper notes 已就绪)。 **Skip if**: 缺少 `outline/outline.yml` 或 `papers/paper_notes.jsonl`。 **Network**: none. **Guardrail**: bullets-only(NO PROSE);每个 claim 至少 2 个证据来源(或显式说明例外)。
.claude/skills/willoscar-claim-evidence-matrix/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-12 | ✗→✓ | ▲ Improved | 12% | 0% |
| case-10 | ✗→✓ | ▲ Improved | -49% | 0% |
| case-11 | ✗→✓ | ▲ Improved | -46% | 0% |
| case-15 | ✗→✓ | ▲ Improved | 4% | 0% |
| case-18 | ✗→✓ | ▲ Improved | -10% | 0% |
Make the survey’s claims explicit and auditable before writing prose.
This should stay bullets-only (NO PROSE). The goal is to make later writing easy and to prevent “template prose” from sneaking in.
outline/outline.ymlpapers/paper_notes.jsonloutline/mapping.tsvoutline/claim_evidence_matrix.mdUses: outline/outline.yml, outline/mapping.tsv.
bibkey exists in papers/paper_notes.jsonl, include [@BibKey] next to evidence items to make later prose/LaTeX conversion smoother.uv run python .codex/skills/claim-evidence-matrix/scripts/run.py --helpuv run python .codex/skills/claim-evidence-matrix/scripts/run.py --workspace <workspace>--help (this helper is intentionally minimal)outline/claim_evidence_matrix.md by tightening claims and adding caveats when evidence is abstract-level.pipeline.py --strict it will be blocked only if placeholder markers remain.Fix:
[@BibKey] because keys are missingFix:
citation-verifier to generate citations/ref.bib, then use the produced keys in the matrix.| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 4,367 | 4,462 | +2% | 1 | 1 | 0% | 740 | 781 | +6% | 0 | 0 | — |
case-02 | fail→fail | 25,472 | 5,182 | -80% | 1 | 1 | 0% | 4,291 | 877 | -80% | 0 | 0 | — |
case-12 | fail→pass | 4,920 | 1,713 | -65% | 1 | 1 | 0% | 765 | 854 | +12% | 0 | 0 | — |
case-03 | fail→fail | 5,956 | 5,346 | -10% | 1 | 1 | 0% | 925 | 837 | -10% | 0 | 0 | — |
case-04 | fail→fail | 4,469 | 5,831 | +30% | 1 | 1 | 0% | 687 | 872 | +27% | 0 | 0 | — |
case-05 | fail→fail | 4,119 | 6,916 | +68% | 1 | 1 | 0% | 730 | 1,011 | +38% | 0 | 0 | — |
case-06 | fail→fail | 2,156 | 5,516 | +156% | 1 | 1 | 0% | 350 | 985 | +181% | 0 | 0 | — |
case-07 | fail→fail | 4,553 | 5,207 | +14% | 1 | 1 | 0% | 729 | 883 | +21% | 0 | 0 | — |
case-08 | fail→fail | 10,433 | 3,422 | -67% | 1 | 1 | 0% | 1,728 | 1,161 | -33% | 0 | 0 | — |
case-09 | pass→pass | 9,715 | 4,012 | -59% | 1 | 1 | 0% | 1,627 | 1,260 | -23% | 0 | 0 | — |
case-10 | fail→pass | 12,222 | 3,852 | -68% | 1 | 1 | 0% | 2,507 | 1,267 | -49% | 0 | 0 | — |
case-11 | fail→pass | 11,860 | 2,906 | -75% | 1 | 1 | 0% | 1,952 | 1,062 | -46% | 0 | 0 | — |
case-13 | pass→pass | 16,888 | 2,125 | -87% | 1 | 1 | 0% | 1,475 | 922 | -37% | 0 | 0 | — |
case-14 | pass→pass | 9,194 | 4,788 | -48% | 1 | 1 | 0% | 1,648 | 1,376 | -17% | 0 | 0 | — |
case-15 | fail→pass | 9,114 | 4,682 | -49% | 1 | 1 | 0% | 1,287 | 1,343 | +4% | 0 | 0 | — |
case-16 | pass→pass | 8,236 | 4,346 | -47% | 1 | 1 | 0% | 1,430 | 1,400 | -2% | 0 | 0 | — |
case-17 | pass→pass | 12,045 | 5,024 | -58% | 1 | 1 | 0% | 1,798 | 1,347 | -25% | 0 | 0 | — |
case-18 | fail→pass | 12,803 | 6,340 | -50% | 1 | 1 | 0% | 1,786 | 1,603 | -10% | 0 | 0 | — |
case-19 | fail→pass | 12,091 | 4,135 | -66% | 1 | 1 | 0% | 1,781 | 1,251 | -30% | 0 | 0 | — |
case-20 | fail→pass | 3,200 | 30,245 | +845% | 1 | 1 | 0% | 466 | 5,199 | +1016% | 0 | 0 | — |
case-21 | pass→fail | 12,922 | 4,881 | -62% | 1 | 1 | 0% | 2,435 | 774 | -68% | 0 | 0 | — |
case-22 | fail→fail | 6,007 | 6,238 | +4% | 1 | 1 | 0% | 890 | 970 | +9% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 13 counted toward the lift figure. The other 9 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +27 percentage points is the difference between those two pass rates over the 13 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.