Install any skill in seconds. Free to start, no credit card required.
Get Started Free →学术参考文献的端到端检查、验证与修复。解析原始文献列表,去重,按APA第七版(APA 7th edition)格式化,通过OpenAlex和Semantic Scholar数据库验证每条文献的准确性,自动修复DOI、年份、标题等问题,并生成人工审核清单和Word文档。Use this skill whenever you need to: 参考文献检查, 参考文献格式调整, APA格式化, 文献去重, DOI验证, 参考文献修复, 参考文献整理, reference check, APA formatting, bibliography verification, DOI validation, citation formatting, or reference list cleanup.
.claude/skills/o0000-code-academic-ref-check/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-06 | ✗→✓ | ▲ Improved | 120% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 19% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 163% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 118% | 0% |
| case-14 | ✗→✓ | ▲ Improved | 95% | 0% |
学术参考文献的完整检查、验证、修复、格式化流程。从原始文献列表到格式规范、数据库验证通过、附带人工审核清单的终版文献输出。
以下五条规则是整个流程的基石。违反任何一条都会导致下游产出错误,且错误难以被后续阶段发现。
所有阶段引用文献时使用内容指纹(作者姓_年份_标题关键词),不使用位置序号(如中[3]、英[218])。这是因为去重、修复、重排序都会改变文献位置,位置序号在第一次修改后就会指向错误的文献,而指纹ID与文献内容绑定,无论如何重排都保持有效。
指纹ID生成规则:
{第一作者姓}_{年份}_{标题前3个实词}(跳过a/an/the/of/in/on/for/and/with/to){第一作者姓}_{年份}_{标题前5个字}Harding_2025_Musical_neurodynamics、安心_2017_高海拔驻留时间对学术引用必须精确——不能基于猜测修改文献内容。自动修复仅限于数据库置信度高且修复方向无歧义的场景(如DOI补充、明显年份错误)。涉及作者增删、标题实质性差异等情况,只报告不修改,留给用户判断。具体判断标准见 references/fix_decision_matrix.md。
人工审核清单(或任何面向用户的文档)中引用的"当前内容"必须从终版文献文件中逐字复制。中间阶段的验证报告只用于确定"哪些条目需要关注",不用于提供条目的具体内容。这条规则存在的原因是:验证阶段记录的文献快照会在后续修复中过时,如果清单从旧快照中复制内容,会出现清单描述与终版文件完全不一致的严重错误。
Sentence case、&号、DOI格式等规则性问题由全量扫描式SubAgent逐条检查,确保100%覆盖。期刊真实性、引用准确性等需要判断力的问题由领域专家评审。两者并行运行但职责不重叠。混合使用会导致规则性问题被专家式抽检遗漏(首次执行中58条sentence case问题被4位专家全部漏掉,就是这个教训)。
人工审核清单必须在所有自动修复和专家评审完成后、作为最终阶段独立生成。过早生成会包含已被后续修复的问题,导致清单与终版文件不一致。
每个SubAgent加载本Skill后,根据自己被分配的角色,阅读对应的操作指南和reference文件。
职责: 将原始文献文本解析为结构化数据,为每条文献分配指纹ID。
操作要点:
[2]、中文[J]标注等)去重规则:
职责: 将结构化文献数据按APA第七版格式输出。
Before formatting, read references/apa7_rules.md for the complete APA 7th edition formatting rules.
操作要点:
姓, 名首字母.,1-20位全列,21+使用省略号规则*期刊名*标记)https://doi.org/前缀格式&连接最后两位作者职责: 通过学术数据库验证每条文献的准确性。
Before verifying, read references/verification_guide.md for the complete database query workflow and result classification rules.
核验主路径 = scripts/verify_http.py(OpenAlex + CrossRef 公开 HTTP API,无需 key)。 若运行环境未提供 semantic-scholar / openalex MCP server(常见默认情况),原"强制加载 SS/OpenAlex MCP"路径会空转。本 Skill 接受这一降级,核验改走 HTTP 脚本:
bashpython3 scripts/verify_http.py --in refs.json --out-dir <dir> # 结构化 JSON 输入 python3 scripts/verify_http.py --in-text refs.md --out-dir <dir> # 原始文本/Markdown 引用列表(自解析)
L11_ref_verify_report.md(人读,绿/黄/红/unverified 四态)+ L11_ref_verify_report.json(机读)。MCP 通道(可选 · 仅当环境中存在 SS/OpenAlex MCP 时): 若所在环境确实装有 semantic-scholar / openalex MCP server,可将其作为补充核验通道与 HTTP 脚本交叉印证(用 ToolSearch 加载 mcp__semantic-scholar__* / mcp__openalex__* 后调用,工具用法见 references/verification_guide.md)。这是条件分支,不是前置必需步骤——若环境不具备这些 MCP,直接走 HTTP 主路径即可,无需尝试加载。
验证策略: 主路径下由 verify_http.py 内部完成「DOI 优先 → 标题搜索 → 自算相似度判定」并按绿/黄/红/unverified 分类(见下方 ABCD 类对照);MCP 在场时可对黄/红条目再交叉印证。两库(HTTP 或 MCP)都查不到才标记为 D 类。
结果分类: | 类别 | 含义 | 后续处理 | |------|------|---------| | A | 验证通过 | 无需处理 | | B | 可自动修复(置信度高) | 交给修复器 | | C | 需人工判断 | 列入审核清单 | | D | 未查到 | 根据文献类型决定是否列入清单 |
中文文献特殊处理: OpenAlex和Semantic Scholar对中文文献覆盖有限。中文文献查不到是正常的,不应被视为信息可能有误的信号,除非文献本身存在格式异常。
职责: 汇总验证结果,对B类条目执行自动修复。
Before fixing, read references/fix_decision_matrix.md for the auto-fix vs report-only decision rules.
操作要点:
典型可自动修复的场景: DOI补充/格式修正、明显年份错误(如2026→2025)、sentence case转换、期刊名PubMed注释去除、标点符号修正、排序修正。
典型仅报告的场景: 作者列表大幅变动、标题实质性差异、卷期页码大幅不一致、数据库未收录的文献。
职责: 对终版文献执行全量逐条规则检查,确保100%覆盖。
Before scanning, read references/apa7_rules.md for the complete checklist of rules to verify.
与专家评审的区别: 规则扫描器检查的是可程序化判断的格式规则(有明确的对/错标准),专家评审处理的是需要学术判断力的问题。
必须逐条扫描的规则(每条文献都检查):
&号(中英文文献均需检查)https://doi.org/前缀,无多余空格/重复前缀)*期刊名*)职责: 执行需要学术判断力的非规则性审查。
审查维度:
不负责的事项: sentence case、&号等规则性检查(由规则扫描器负责)。
职责: 生成面向用户的人工审核清单。这是整个流程中对准确性要求最高的环节。
Before generating, read references/checklist_spec.md for the complete checklist format specification, category definitions, and validation rules.
强制执行的生成流程:
scripts/validate_checklist.py)五类分类:
职责: 将终版Markdown文献文件转换为符合博士论文排版标准的Word文档。
使用 scripts/convert_refs.py 执行转换:
bashpython scripts/convert_refs.py <input.md> [output.docx]
格式规范:
*期刊名*标记自动识别)依赖: pip install python-docx
以下是最常用的10条规则,完整规则见 references/apa7_rules.md。
| # | 规则 | 正确示例 | |---|------|---------| | 1 | 作者格式:姓, 名首字母. | Zhang, L. M. | | 2 | 多作者用,分隔,最后两位用& | Li, A., Wang, B., & Chen, C. | | 3 | 21+作者:前19位...最后1位 | Author, A., Author, B., ... Author, U. | | 4 | 年份在作者后括号内 | Smith, J. (2023). | | 5 | 文章标题sentence case | Effects of music on cognitive development in children | | 6 | 期刊名title case+斜体 | *Journal of Experimental Psychology* | | 7 | 卷号斜体,期号不斜体括号内 | *12*(3), 45-67 | | 8 | DOI用https://doi.org/前缀 | https://doi.org/10.1037/rev0000106 | | 9 | 书名斜体+sentence case | *Cognitive psychology: A student's handbook* | | 10 | 条目末尾无句号(如以DOI结尾) | DOI链接后不加句号 |
中文文献补充规则:
&连接最后两位作者(不用"和")作者. (年份). *标题* [博士学位论文, 院校名称]. 数据库名称.bash# 基本用法 python scripts/convert_refs.py input.md output.docx # 默认输出(与输入同名.docx) python scripts/convert_refs.py input.md
脚本自动处理:元数据跳过、分类标题识别、期刊名斜体渲染、方括号注释清理。
bashpython scripts/validate_checklist.py checklist.md final_refs.md
执行五项校验:
scripts/verify_http.py)/ MCP 在场时可选补充——运行环境无 SS/OpenAlex MCP,接受降级走 OpenAlex/CrossRef 公开 HTTP API。清除原强制加载 MCP 块中"环境中 MCP 一定可加载"的过度断言与"必须先加载 MCP 才能核验"的前置强制,并移除 S3 留下的本节待办标注。verification_guide.md §1 同步收口。| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 11,899 | 6,442 | -46% | 1 | 1 | 0% | 1,868 | 4,830 | +159% | 0 | 0 | — |
case-02 | fail→fail | 11,437 | 6,915 | -40% | 1 | 1 | 0% | 1,811 | 4,896 | +170% | 0 | 0 | — |
case-03 | fail→fail | 12,424 | 10,197 | -18% | 1 | 1 | 0% | 2,031 | 5,373 | +165% | 0 | 0 | — |
case-04 | fail→fail | 14,984 | 18,233 | +22% | 1 | 1 | 0% | 2,738 | 7,133 | +161% | 0 | 0 | — |
case-05 | fail→fail | 15,657 | 16,159 | +3% | 1 | 1 | 0% | 2,751 | 6,562 | +139% | 0 | 0 | — |
case-06 | fail→pass | 14,738 | 9,639 | -35% | 1 | 1 | 0% | 2,538 | 5,589 | +120% | 0 | 0 | — |
case-07 | fail→fail | 8,668 | 5,074 | -41% | 1 | 1 | 0% | 1,459 | 4,650 | +219% | 0 | 0 | — |
case-08 | pass→pass | 13,915 | 6,045 | -57% | 1 | 1 | 0% | 2,477 | 4,795 | +94% | 0 | 0 | — |
case-09 | fail→pass | 20,718 | 6,330 | -69% | 1 | 1 | 0% | 4,155 | 4,950 | +19% | 0 | 0 | — |
case-10 | pass→pass | 17,305 | 3,451 | -80% | 1 | 1 | 0% | 1,289 | 4,394 | +241% | 0 | 0 | — |
case-11 | pass→pass | 14,505 | 6,442 | -56% | 1 | 1 | 0% | 2,222 | 4,782 | +115% | 0 | 0 | — |
case-12 | fail→pass | 13,134 | 9,490 | -28% | 1 | 1 | 0% | 2,014 | 5,291 | +163% | 0 | 0 | — |
case-13 | fail→pass | 11,798 | 6,684 | -43% | 1 | 1 | 0% | 2,215 | 4,827 | +118% | 0 | 0 | — |
case-14 | fail→pass | 14,422 | 6,210 | -57% | 1 | 1 | 0% | 2,479 | 4,834 | +95% | 0 | 0 | — |
case-15 | fail→pass | 12,058 | 4,557 | -62% | 1 | 1 | 0% | 2,230 | 4,597 | +106% | 0 | 0 | — |
case-16 | pass→pass | 8,216 | 6,963 | -15% | 1 | 1 | 0% | 1,533 | 5,022 | +228% | 0 | 0 | — |
case-17 | pass→pass | 8,065 | 3,809 | -53% | 1 | 1 | 0% | 1,465 | 4,426 | +202% | 0 | 0 | — |
case-18 | pass→pass | 21,575 | 9,418 | -56% | 1 | 1 | 0% | 2,251 | 5,232 | +132% | 0 | 0 | — |
case-19 | pass→pass | 11,512 | 6,987 | -39% | 1 | 1 | 0% | 2,219 | 4,994 | +125% | 0 | 0 | — |
case-20 | fail→pass | 9,244 | 3,192 | -65% | 1 | 1 | 0% | 1,623 | 4,335 | +167% | 0 | 0 | — |
case-21 | fail→pass | 12,261 | 3,844 | -69% | 1 | 1 | 0% | 2,354 | 4,501 | +91% | 0 | 0 | — |
case-22 | fail→pass | 9,201 | 2,302 | -75% | 1 | 1 | 0% | 1,533 | 4,183 | +173% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +41 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.