Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Validate bibliography entries against citations in all lecture files. Structural checks (missing/unused entries, malformed fields) by default; `--semantic` adds citation-drift detection, DOI verification, and style-consistency checks.
.claude/skills/pedrohcgs-validate-bib/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | 834% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 42% | 0% |
| case-09 | ✗→✓ | ▲ Improved | -17% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 25% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 32% | 0% |
Cross-reference citations in lecture files against bibliography entries. Two modes:
--semantic: adds citation-drift detection (duplicate entries for the same paper), DOI verification via crossref, and citation-style consistency within each file.Report saved to quality_reports/bib_audit_[structural|semantic].md.
.tex: \cite{, \citet{, \citep{, \citeauthor{, \citeyear{, \textcite{, \parencite{.qmd / .md: @key, [@key], [@key1; @key2].bib..bib but never cited..bib key (e.g., Smith2020 vs Smth2020).doi field normalized (no leading https://doi.org/).quality_reports/bib_audit_structural.md.Slides/*.tex
Quarto/*.qmd
guide/*.qmd
master_supporting_docs/**/*.texBibliography_base.bib at repo root by default; override via CLAUDE.md.
--semantic)Everything in Mode 1, plus:
Multiple .bib entries describing the same paper under different keys. Symptoms:
Smith2020 + Smith2020a with identical DOI or title.CallawaySantAnna2021 + CS2021 both pointing to the same paper..bib files.Detection heuristics (any → FLAG):
| Check | Signal | |---|---| | Same DOI across keys | Hard-duplicate (CRITICAL) | | Same title (case-insensitive, punct-stripped) | Likely duplicate (CRITICAL) | | Same author+year+journal | Probable duplicate (MEDIUM) | | Title Jaccard > 0.85 on tokens ≥ 4 chars | Soft-duplicate (LOW) |
For each flagged pair: list both keys, where each is cited, and recommend a canonical key (prefer most-cited, then alphabetically first).
For each entry with a doi, fetch https://api.crossref.org/works/{doi} and compare:
Severity:
Rate limit: cap 50 lookups per run, 0.5s delay between calls. Cache in quality_reports/.doi_cache.json.
Opt-out: --skip-doi for offline or no-WebFetch environments.
For each file, count citation commands (\citet vs \citep vs \cite; @key vs [@key]). FLAG files with mixed styles without an obvious pattern (e.g., 20× \citep and 3× \cite in the same deck). Low-severity.
Gated behind --cite-claim. For the top-10 most-cited works per file, WebFetch the crossref abstract and surface it beside the in-text context. No auto-judgment — humans decide if the claim matches.
> This is existence/structure, not appropriateness. Deciding whether the cited paper actually says what the in-text claim attributes to it is /verify-claims's job — it reads the source and grounds a supports / partial / contradicts verdict in quotes + pages (with the EXPLAINED escape for a defensible named alternative). --cite-claim only surfaces the abstract; for the verdict, run /verify-claims.
quality_reports/bib_audit_semantic.md)markdown# Bibliography Semantic Audit **Date:** YYYY-MM-DD **Bibliography:** Bibliography_base.bib (N entries) **Files scanned:** [list] ## Summary | Check | Critical | Medium | Low | |---|---|---|---| | Structural | | | | | Citation drift | | | | | DOI verification | | | | | Style consistency | 0 | 0 | | ## Critical Issues ### Duplicate entries | Keys | Signal | Citations | Recommended canonical | |---|---|---|---| ### DOI mismatches | Key | Field | .bib value | crossref value | |---|---|---|---| ## Medium / Low issues … ## Next steps 1. Resolve duplicates — pick canonical key, update citations, remove orphans. 2. Fix DOI mismatches — verify paper in crossref or strip the wrong DOI. 3. Review style-consistency notes.
.claude/skills/review-paper/SKILL.md — pair for full pre-submission..claude/skills/audit-reproducibility/SKILL.md — numeric-claims counterpart..claude/skills/verify-claims/SKILL.md — citation appropriateness counterpart (does the cited paper support the claim?). This skill checks that a citation exists and is well-formed; /verify-claims checks that it holds./verify-claims's job (see 2d); this skill stays existence-and-structure only..bib file — all edits are recommendations.| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 21,993 | 4,799 | -78% | 1 | 1 | 0% | 4,745 | 1,886 | -60% | 0 | 0 | — |
case-02 | fail→fail | 30,403 | 6,808 | -78% | 1 | 1 | 0% | 6,195 | 2,034 | -67% | 0 | 0 | — |
case-03 | fail→fail | 4,190 | 4,033 | -4% | 1 | 1 | 0% | 714 | 1,835 | +157% | 0 | 0 | — |
case-04 | fail→pass | 2,278 | 10,470 | +360% | 1 | 1 | 0% | 376 | 3,511 | +834% | 0 | 0 | — |
case-05 | fail→fail | 16,117 | 5,253 | -67% | 1 | 1 | 0% | 3,055 | 2,229 | -27% | 0 | 0 | — |
case-06 | fail→fail | 2,679 | 3,798 | +42% | 1 | 1 | 0% | 430 | 1,801 | +319% | 0 | 0 | — |
case-07 | pass→pass | 12,821 | 11,206 | -13% | 1 | 1 | 0% | 2,239 | 3,508 | +57% | 0 | 0 | — |
case-12 | fail→pass | 7,482 | 1,612 | -78% | 1 | 1 | 0% | 1,311 | 1,857 | +42% | 0 | 0 | — |
case-08 | pass→pass | 10,734 | 4,043 | -62% | 1 | 1 | 0% | 1,938 | 2,236 | +15% | 0 | 0 | — |
case-09 | fail→pass | 13,811 | 2,954 | -79% | 1 | 1 | 0% | 2,420 | 2,002 | -17% | 0 | 0 | — |
case-10 | fail→pass | 10,044 | 3,357 | -67% | 1 | 1 | 0% | 1,641 | 2,055 | +25% | 0 | 0 | — |
case-11 | fail→pass | 8,731 | 1,730 | -80% | 1 | 1 | 0% | 1,429 | 1,888 | +32% | 0 | 0 | — |
case-13 | pass→pass | 7,125 | 1,572 | -78% | 1 | 1 | 0% | 1,294 | 1,799 | +39% | 0 | 0 | — |
case-14 | fail→fail | 11,561 | 1,621 | -86% | 1 | 1 | 0% | 1,810 | 1,817 | +0% | 0 | 0 | — |
case-15 | pass→pass | 10,846 | 2,046 | -81% | 1 | 1 | 0% | 1,649 | 1,898 | +15% | 0 | 0 | — |
case-16 | fail→pass | 12,567 | 2,744 | -78% | 1 | 1 | 0% | 2,045 | 2,001 | -2% | 0 | 0 | — |
case-21 | fail→pass | 11,790 | 1,853 | -84% | 1 | 1 | 0% | 1,941 | 1,874 | -3% | 0 | 0 | — |
case-17 | fail→pass | 11,471 | 1,552 | -86% | 1 | 1 | 0% | 2,003 | 1,804 | -10% | 0 | 0 | — |
case-18 | pass→pass | 15,040 | 4,402 | -71% | 1 | 1 | 0% | 2,489 | 2,275 | -9% | 0 | 0 | — |
case-19 | fail→pass | 10,273 | 2,887 | -72% | 1 | 1 | 0% | 1,726 | 2,056 | +19% | 0 | 0 | — |
case-20 | fail→pass | 9,144 | 1,325 | -86% | 1 | 1 | 0% | 1,498 | 1,792 | +20% | 0 | 0 | — |
case-22 | pass→pass | 9,110 | 1,777 | -80% | 1 | 1 | 0% | 1,346 | 1,811 | +35% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 18 counted toward the lift figure. The other 4 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +45 percentage points is the difference between those two pass rates over the 18 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.