Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when adding, verifying or cleaning citations and BibTeX entries in Stage 07 (Writing), when a reference cannot be resolved cleanly from DBLP or CrossRef, when checking that a cited paper actually supports the claim attributed to it, or when filling citation_verification.json.
.claude/skills/tangxiangru-citation-discipline/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-10 | ✗→✓ | ▲ Improved | -8% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 22% | 0% |
| case-13 | ✗→✓ | ▲ Improved | -20% | 0% |
| case-04 | ✓→✓ | = Same ✓ | 22% | 0% |
| case-05 | ✓→✓ | = Same ✓ | 63% | 0% |
Stage 07 must write workspace/artifacts/citation_verification.json, and the gate checks it has a non-empty claim_coverage list where every entry carries a claim and at least one citation_keys or source_ids value. That file is a record of verification, so it is only worth anything if verification happened.
reference.md in this directory is the long-form treatment: source APIs, the full verification workflow, BibTeX field rules, entry templates by publication type, and a troubleshooting table. Read it when a specific reference will not resolve.
Never generate a citation from memory.
The dangerous failure is not an obviously fake reference. It is one that looks plausible: real authors with a fabricated title, a real title with the wrong year, a real arXiv ID attached to the wrong venue, a preprint and its published version silently merged into one entry. These survive a read-through and surface during review.
If a citation cannot be verified programmatically or from the run's own workspace/literature/sources.json, mark it unresolved and say so. Do not produce a plausible-looking BibTeX entry to fill the hole.
For each citation:
CrossRef for DOIs, the publisher page, or arXiv for preprints.
any one of them means you have the wrong record, not a typo to smooth over.
are attributing to it. A correctly-formatted citation for a claim the paper does not make is still a fabricated citation.
citation_verification.json with the claim it supports.workspace/literature/sources.json and claims.json were built in Stage 01 with source IDs. A Stage 07 claim that traces back to a Stage 01 claim should reuse its source_id rather than re-deriving a citation. If a Stage 07 claim has no Stage 01 ancestor, that is worth noticing: it may be a claim the run never gathered evidence for.
cite it as a preprint rather than dressing it as a conference paper.
author2024shorttitle), never renumbered.working bibliography to BibLaTeX mid-run.
\cite key resolves to an entry in the .bib..bib is cited at least once.citation_verification.json covers each substantive claim, not just theeasy ones.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 14,535 | 14,430 | -1% | 1 | 1 | 0% | 204 | 923 | +352% | 0 | 0 | — |
case-02 | fail→fail | 15,216 | 9,489 | -38% | 1 | 1 | 0% | 359 | 946 | +164% | 0 | 0 | — |
case-03 | fail→fail | 14,721 | 6,549 | -56% | 1 | 1 | 0% | 248 | 927 | +274% | 0 | 0 | — |
case-04 | pass→pass | 13,303 | 9,590 | -28% | 1 | 1 | 0% | 1,226 | 1,496 | +22% | 0 | 0 | — |
case-05 | pass→pass | 5,236 | 9,205 | +76% | 1 | 1 | 0% | 916 | 1,489 | +63% | 0 | 0 | — |
case-06 | fail→fail | 10,518 | 12,770 | +21% | 1 | 1 | 0% | 1,895 | 2,127 | +12% | 0 | 0 | — |
case-07 | pass→pass | 16,251 | 8,100 | -50% | 1 | 1 | 0% | 2,777 | 2,105 | -24% | 0 | 0 | — |
case-08 | pass→pass | 15,691 | 12,656 | -19% | 1 | 1 | 0% | 1,726 | 1,928 | +12% | 0 | 0 | — |
case-09 | pass→pass | 15,895 | 10,832 | -32% | 1 | 1 | 0% | 1,895 | 1,519 | -20% | 0 | 0 | — |
case-10 | fail→pass | 15,314 | 9,309 | -39% | 1 | 1 | 0% | 1,559 | 1,440 | -8% | 0 | 0 | — |
case-11 | fail→pass | 13,961 | 11,115 | -20% | 1 | 1 | 0% | 1,454 | 1,768 | +22% | 0 | 0 | — |
case-12 | pass→pass | 13,005 | 9,695 | -25% | 1 | 1 | 0% | 1,289 | 1,545 | +20% | 0 | 0 | — |
case-13 | fail→pass | 10,139 | 4,019 | -60% | 1 | 1 | 0% | 1,729 | 1,379 | -20% | 0 | 0 | — |
case-14 | pass→pass | 7,719 | 7,819 | +1% | 1 | 1 | 0% | 1,306 | 1,106 | -15% | 0 | 0 | — |
case-15 | pass→pass | 9,430 | 9,554 | +1% | 1 | 1 | 0% | 1,620 | 1,416 | -13% | 0 | 0 | — |
case-16 | pass→pass | 22,291 | 3,292 | -85% | 1 | 1 | 0% | 2,947 | 1,309 | -56% | 0 | 0 | — |
case-17 | pass→pass | 15,672 | 8,314 | -47% | 1 | 1 | 0% | 1,443 | 1,306 | -9% | 0 | 0 | — |
case-18 | pass→pass | 14,952 | 13,058 | -13% | 1 | 1 | 0% | 1,768 | 2,237 | +27% | 0 | 0 | — |
case-19 | pass→pass | 13,047 | 2,809 | -78% | 1 | 1 | 0% | 1,275 | 1,095 | -14% | 0 | 0 | — |
case-20 | pass→pass | 16,016 | 9,327 | -42% | 1 | 1 | 0% | 1,740 | 1,445 | -17% | 0 | 0 | — |
case-21 | pass→pass | 12,599 | 13,148 | +4% | 1 | 1 | 0% | 1,408 | 2,204 | +57% | 0 | 0 | — |
case-22 | pass→pass | 21,681 | 23,376 | +8% | 1 | 1 | 0% | 2,765 | 3,680 | +33% | 0 | 0 | — |
case-23 | pass→pass | 12,070 | 16,162 | +34% | 1 | 1 | 0% | 2,083 | 3,100 | +49% | 0 | 0 | — |
case-24 | pass→pass | 18,690 | 12,095 | -35% | 1 | 1 | 0% | 2,327 | 2,631 | +13% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 24 cases were attempted, and 21 counted toward the lift figure. The other 3 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +13 percentage points is the difference between those two pass rates over the 21 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.