Install any skill in seconds. Free to start, no credit card required.
Get Started Free →用于终稿完成且脚注需要后处理时:去重 [^key] 引用,转换为 [N] 编号,并追加参考文献。
.claude/skills/opensensenova-sn-prepare-citations/SKILL.md| Model | Eval pass | Runs |
|---|---|---|
| gemini-3.6-flash | 100% | 29 |
| gemini-3.1-pro-preview | 100% | 1 |
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 312% | 0% |
| case-04 | ✗→✓ | ▲ Improved | -15% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 12% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 6% | 0% |
| case-14 | ✗→✓ | ▲ Improved | -16% | 0% |
处理 sn-deep-research 的 stitched.md:从 evidence.json 收集 sources[],将正文中的 [^source_id] 转换为 [N] 编号引用,插入 L0/TOC,并追加参考文献。
bashpython3 scripts/prepare_citations.py \ --report <report_dir>/stitched.md \ --evidence <report_dir>/sub_reports/d1.evidence.json <report_dir>/sub_reports/d2.evidence.json \ --outline <report_dir>/outline.json \ --output <report_dir>/report.md
| 参数 | 说明 | |---|---| | --report | 输入 markdown。sn-deep-research 中通常是 stitched.md | | --evidence | 全部 d*.evidence.json,用于收集 source 元数据和修复 claim-id 泄漏 | | --outline | 可选但推荐。提供 L0、TOC 和标题结构信息 | | --output | 输出 markdown。sn-deep-research 中通常是 report.md | | --no-l0 | 关闭 L0 摘要层渲染 | | --no-toc | 关闭 TOC 渲染 |
sources[] 收集引用元数据。[^dN.cM] claim-id 引用泄漏;能映射到 claim evidence 时替换为对应 source_id,不能映射则报告 unresolved。[^source_id],按首次出现顺序分配编号。[^source_id] → [N],移除脚注定义行。## 参考文献。report.md 和同目录 citations.json。orphan_citations 非空:不要交付,回 writer/stitcher 修正引用或删除 unsupported 内容。claim_id_leakage.unresolved 非空:不要交付,回 writer 修正 [^dN.cM]。claim_id_leakage.resolved 非空但 unresolved 为空:可以继续,但记录为警告;writer 后续应直接输出 [^source_id]。不传 --outline / --output 时,脚本会覆写 --report 指向的文件,仅做引用编号和参考文献追加。sn-deep-research 正常流程不使用该模式。
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-10 | pass→pass | 8,395 | 5,639 | -33% | 1 | 1 | 0% | 1,486 | 1,667 | +12% | 0 | 0 | — |
case-01 | fail→pass | 6,595 | 2,053 | -69% | 1 | 1 | 0% | 276 | 1,138 | +312% | 0 | 0 | — |
case-02 | fail→fail | 3,695 | 4,875 | +32% | 1 | 1 | 0% | 241 | 986 | +309% | 0 | 0 | — |
case-03 | fail→fail | 4,517 | 4,990 | +10% | 1 | 1 | 0% | 324 | 1,058 | +227% | 0 | 0 | — |
case-04 | fail→pass | 8,122 | 3,182 | -61% | 1 | 1 | 0% | 1,431 | 1,210 | -15% | 0 | 0 | — |
case-05 | pass→pass | 11,234 | 4,353 | -61% | 1 | 1 | 0% | 1,964 | 1,491 | -24% | 0 | 0 | — |
case-06 | pass→pass | 8,740 | 3,810 | -56% | 1 | 1 | 0% | 1,513 | 1,362 | -10% | 0 | 0 | — |
case-07 | pass→pass | 3,829 | 3,591 | -6% | 1 | 1 | 0% | 748 | 1,373 | +84% | 0 | 0 | — |
case-08 | fail→pass | 7,758 | 4,592 | -41% | 1 | 1 | 0% | 1,351 | 1,519 | +12% | 0 | 0 | — |
case-09 | fail→pass | 6,000 | 2,550 | -57% | 1 | 1 | 0% | 1,075 | 1,144 | +6% | 0 | 0 | — |
case-11 | fail→fail | 9,831 | 1,459 | -85% | 1 | 1 | 0% | 1,642 | 872 | -47% | 0 | 0 | — |
case-12 | fail→fail | 3,083 | 1,369 | -56% | 1 | 1 | 0% | 451 | 853 | +89% | 0 | 0 | — |
case-13 | pass→pass | 10,298 | 3,363 | -67% | 1 | 1 | 0% | 1,839 | 1,228 | -33% | 0 | 0 | — |
case-14 | fail→pass | 8,692 | 3,361 | -61% | 1 | 1 | 0% | 1,513 | 1,265 | -16% | 0 | 0 | — |
case-15 | pass→pass | 8,540 | 5,559 | -35% | 1 | 1 | 0% | 1,525 | 1,703 | +12% | 0 | 0 | — |
case-16 | fail→pass | 10,152 | 3,827 | -62% | 1 | 1 | 0% | 2,032 | 1,419 | -30% | 0 | 0 | — |
case-17 | pass→pass | 7,040 | 3,399 | -52% | 1 | 1 | 0% | 1,410 | 1,351 | -4% | 0 | 0 | — |
case-18 | fail→pass | 7,612 | 2,870 | -62% | 1 | 1 | 0% | 1,433 | 1,240 | -13% | 0 | 0 | — |
case-19 | pass→pass | 5,434 | 2,133 | -61% | 1 | 1 | 0% | 939 | 1,055 | +12% | 0 | 0 | — |
case-20 | fail→fail | 20,652 | 17,756 | -14% | 1 | 1 | 0% | 4,054 | 4,073 | +0% | 0 | 0 | — |
case-21 | pass→pass | 5,480 | 7,672 | +40% | 1 | 1 | 0% | 1,105 | 2,203 | +99% | 0 | 0 | — |
case-22 | pass→pass | 10,669 | 10,984 | +3% | 1 | 1 | 0% | 2,127 | 2,906 | +37% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 19 counted toward the lift figure. The other 3 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +32 percentage points is the difference between those two pass rates over the 19 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.