Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Generate and verify BibTeX entries from paper notes, writing `citations/ref.bib` and `citations/verified.jsonl`. **Trigger**: citation, BibTeX, ref.bib, verified.jsonl, references, 引用, 参考文献. **Use when**: 已有 `papers/paper_notes.jsonl`,需要为 prose/LaTeX 准备可追溯的引用(每条都有 url/date/title 验证记录)。
.claude/skills/willoscar-citation-verifier/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-05 | ✗→✓ | ▲ Improved | -12% | 0% |
| case-06 | ✗→✓ | ▲ Improved | -1% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -6% | 0% |
| case-08 | ✗→✓ | ▲ Improved | -32% | 0% |
| case-10 | ✗→✓ | ▲ Improved | -6% | 0% |
Generate citations/ref.bib and ensure every entry has a traceable verification record in citations/verified.jsonl.
When network access is restricted, prefer a “record now, verify later” workflow: keep URLs/titles consistent and leave a clear verification note.
papers/paper_notes.jsonlcitations/ref.bibcitations/verified.jsonlbibkey, title, url, year, authors from papers/paper_notes.jsonl.citations/ref.bib:arxiv_id / primary_category exist (eprint, archivePrefix, primaryClass).citations/verified.jsonl with at least:bibkey, title, url, datenotes field (e.g., “auto-generated; needs manual verification”) and/or request human confirmation depending on your policy.verified.jsonl record.url/date/title in verification records.When network access is restricted, run in offline mode to produce auditable records now, then verify later.
verification_status: offline_generated--verify-onlyverification_statusoffline_generated: record was generated without network verification (needs later verification)verified_online: URL/title verified successfully by the scriptverify_failed: verification was attempted but failed (network error or title mismatch)needs_manual_verification: missing/ambiguous fields (e.g., empty url/title)uv run python .codex/skills/citation-verifier/scripts/run.py --helpuv run python .codex/skills/citation-verifier/scripts/run.py --workspace <workspace> --offline--offline: do not attempt network verification; write verification_status=offline_generated--verify-only: verify existing citations/verified.jsonl records (does not rewrite BibTeX)--verification-note <text>: stored in citations/verified.jsonl notesuv run python .codex/skills/citation-verifier/scripts/run.py --workspace <workspace> --offline --verification-note "auto-generated; needs manual verification"uv run python .codex/skills/citation-verifier/scripts/run.py --workspace <workspace> --verify-onlyurl, date, title.{} in titles to keep bibtex parsing robust.& % $ # _) and rewrites superscript patterns like X^N or X$^N$ as X\textsuperscript{N} to keep LaTeX builds stable.url fields (BibTeX styles wrap them with \url{...}); @misc uses howpublished=\url{...}.offline_generated as a to-do for human/network verification.bibkey / missing url in notesSymptom:
citations/ref.bib is missing entries, or verified.jsonl has empty url/title.Causes:
papers/paper_notes.jsonl lacks bibkey/url fields.Solutions:
bibkey and a canonical url.verification_status=offline_generatedSymptom:
Causes:
--offline was used, or network verification was unavailable.Solutions:
--verify-only to upgrade records.citations/verified.jsonl with notes.citations/verified.jsonl record.url, date, title.| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 5,773 | 4,808 | -17% | 1 | 1 | 0% | 462 | 1,388 | +200% | 0 | 0 | — |
case-02 | fail→fail | 9,717 | 5,865 | -40% | 1 | 1 | 0% | 733 | 1,360 | +86% | 0 | 0 | — |
case-03 | fail→fail | 5,229 | 4,552 | -13% | 1 | 1 | 0% | 266 | 1,276 | +380% | 0 | 0 | — |
case-04 | pass→pass | 12,955 | 10,092 | -22% | 1 | 1 | 0% | 2,450 | 2,885 | +18% | 0 | 0 | — |
case-05 | fail→pass | 9,683 | 2,149 | -78% | 1 | 1 | 0% | 1,587 | 1,401 | -12% | 0 | 0 | — |
case-06 | fail→pass | 10,066 | 2,773 | -72% | 1 | 1 | 0% | 1,572 | 1,556 | -1% | 0 | 0 | — |
case-07 | fail→pass | 9,097 | 1,826 | -80% | 1 | 1 | 0% | 1,419 | 1,334 | -6% | 0 | 0 | — |
case-08 | fail→pass | 13,062 | 1,526 | -88% | 1 | 1 | 0% | 1,939 | 1,309 | -32% | 0 | 0 | — |
case-09 | pass→pass | 11,019 | 5,095 | -54% | 1 | 1 | 0% | 1,929 | 1,937 | +0% | 0 | 0 | — |
case-10 | fail→pass | 11,706 | 4,166 | -64% | 1 | 1 | 0% | 1,810 | 1,708 | -6% | 0 | 0 | — |
case-11 | pass→pass | 9,846 | 2,751 | -72% | 1 | 1 | 0% | 1,626 | 1,514 | -7% | 0 | 0 | — |
case-12 | fail→pass | 8,002 | 2,372 | -70% | 1 | 1 | 0% | 1,178 | 1,476 | +25% | 0 | 0 | — |
case-13 | pass→pass | 13,358 | 5,186 | -61% | 1 | 1 | 0% | 2,120 | 1,877 | -11% | 0 | 0 | — |
case-14 | fail→pass | 14,883 | 4,738 | -68% | 1 | 1 | 0% | 2,176 | 1,976 | -9% | 0 | 0 | — |
case-15 | fail→pass | 18,218 | 3,158 | -83% | 1 | 1 | 0% | 2,893 | 1,561 | -46% | 0 | 0 | — |
case-16 | pass→pass | 11,146 | 5,961 | -47% | 1 | 1 | 0% | 1,805 | 2,163 | +20% | 0 | 0 | — |
case-17 | fail→pass | 8,039 | 2,787 | -65% | 1 | 1 | 0% | 1,283 | 1,581 | +23% | 0 | 0 | — |
case-18 | fail→pass | 9,386 | 1,646 | -82% | 1 | 1 | 0% | 1,327 | 1,307 | -2% | 0 | 0 | — |
case-19 | pass→pass | 10,055 | 4,790 | -52% | 1 | 1 | 0% | 1,679 | 1,867 | +11% | 0 | 0 | — |
case-20 | pass→pass | 4,381 | 4,009 | -8% | 1 | 1 | 0% | 725 | 1,741 | +140% | 0 | 0 | — |
case-21 | pass→pass | 4,231 | 4,632 | +9% | 1 | 1 | 0% | 680 | 1,898 | +179% | 0 | 0 | — |
case-22 | pass→pass | 6,997 | 5,829 | -17% | 1 | 1 | 0% | 1,207 | 1,999 | +66% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 19 counted toward the lift figure. The other 3 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +45 percentage points is the difference between those two pass rates over the 19 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.