Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when `paper-review` needs a claim-by-claim evidence gap report grounded in an extracted claim ledger. **Trigger**: evidence audit, missing evidence, unsupported claims, 审稿证据审计, 证据缺口. **Use when**: `paper-review` 流程中,需要逐条检查 claim 的证据链、缺 baseline、评测薄弱点。 **Skip if**: 缺少 claims 输入(例如还没有 `output/CLAIMS.md`)。 **Network**: none. **Guardrail**: 只写“缺口/风险/下一步验证”,不要替作者补写论述或引入新主张。
.claude/skills/willoscar-evidence-auditor/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-20 | ✗→✓ | ▲ Improved | -36% | 0% |
| case-03 | ✗→✓ | ▲ Improved | -61% | 0% |
| case-05 | ✗→✓ | ▲ Improved | -65% | 0% |
| case-06 | ✗→✓ | ▲ Improved | -25% | 0% |
| case-21 | ✗→✓ | ▲ Improved | -52% | 0% |
paper-review 流程中,需要逐条检查 claim 的证据链、缺 baseline、评测薄弱点。output/CLAIMS.md)。Transforms a claim ledger into a gap report for paper-review.
output/CLAIMS.mdoutput/MISSING_EVIDENCE.mdoutput/EVIDENCE_AUDIT.jsonl (review-evidence-gap.v1)Each gap block should include:
Each JSONL record must retain claim_id and a stable gap_id so the final review can cite the exact failure rather than paraphrasing an unlocated concern.
scripts/run.py should:
It should not invent new claims or rewrite the manuscript.
claim_id| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 4,316 | 4,336 | +0% | 1 | 1 | 0% | 245 | 434 | +77% | 0 | 0 | — |
case-20 | fail→pass | 12,039 | 5,512 | -54% | 1 | 1 | 0% | 1,811 | 1,150 | -36% | 0 | 0 | — |
case-02 | fail→fail | 6,480 | 3,629 | -44% | 1 | 1 | 0% | 327 | 449 | +37% | 0 | 0 | — |
case-03 | fail→pass | 14,912 | 3,811 | -74% | 1 | 1 | 0% | 2,294 | 895 | -61% | 0 | 0 | — |
case-04 | pass→pass | 24,032 | 4,615 | -81% | 1 | 1 | 0% | 1,615 | 1,043 | -35% | 0 | 0 | — |
case-05 | fail→pass | 11,500 | 2,297 | -80% | 1 | 1 | 0% | 1,756 | 618 | -65% | 0 | 0 | — |
case-06 | fail→pass | 12,990 | 7,841 | -40% | 1 | 1 | 0% | 1,873 | 1,413 | -25% | 0 | 0 | — |
case-21 | fail→pass | 15,449 | 4,672 | -70% | 1 | 1 | 0% | 2,122 | 1,019 | -52% | 0 | 0 | — |
case-07 | pass→pass | 11,146 | 3,175 | -72% | 1 | 1 | 0% | 1,567 | 765 | -51% | 0 | 0 | — |
case-08 | fail→pass | 13,134 | 5,008 | -62% | 1 | 1 | 0% | 2,152 | 1,108 | -49% | 0 | 0 | — |
case-09 | fail→fail | 4,524 | 1,905 | -58% | 1 | 1 | 0% | 661 | 531 | -20% | 0 | 0 | — |
case-10 | fail→fail | 7,448 | 1,603 | -78% | 1 | 1 | 0% | 1,141 | 470 | -59% | 0 | 0 | — |
case-11 | fail→fail | 11,933 | 1,947 | -84% | 1 | 1 | 0% | 1,688 | 523 | -69% | 0 | 0 | — |
case-12 | pass→pass | 11,926 | 6,209 | -48% | 1 | 1 | 0% | 1,834 | 1,107 | -40% | 0 | 0 | — |
case-13 | fail→fail | 8,062 | 5,838 | -28% | 1 | 1 | 0% | 1,224 | 1,221 | -0% | 0 | 0 | — |
case-14 | fail→pass | 9,629 | 2,229 | -77% | 1 | 1 | 0% | 1,674 | 569 | -66% | 0 | 0 | — |
case-15 | fail→fail | 7,376 | 2,794 | -62% | 1 | 1 | 0% | 1,082 | 706 | -35% | 0 | 0 | — |
case-16 | fail→fail | 7,217 | 1,810 | -75% | 1 | 1 | 0% | 1,031 | 488 | -53% | 0 | 0 | — |
case-17 | fail→fail | 5,109 | 2,234 | -56% | 1 | 1 | 0% | 718 | 568 | -21% | 0 | 0 | — |
case-18 | pass→fail | 13,798 | 4,096 | -70% | 1 | 1 | 0% | 1,900 | 897 | -53% | 0 | 0 | — |
case-19 | pass→pass | 13,244 | 4,092 | -69% | 1 | 1 | 0% | 1,914 | 865 | -55% | 0 | 0 | — |
case-22 | fail→pass | 2,445 | 3,972 | +62% | 1 | 1 | 0% | 291 | 796 | +174% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 20 counted toward the lift figure. The other 2 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +32 percentage points is the difference between those two pass rates over the 20 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
The publisher has shipped newer versions since this run, so these numbers describe v1, not the version currently listed.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.