Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Triple-pass verification of every manuscript claim against code, data, refs, and renderer; repair prose while staying renderable. USE WHEN pre-submission, pre-Zenodo, pre-arXiv, abstract numbers disagree with CSV, citations do not support sentences, or user asks to triple-check / verify every claim — even without docs/prompts. Not for casual PDF summary.
.claude/skills/docxology-template-manuscript-claim-verification/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-13 | ✗→✓ | ▲ Improved | -11% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 1% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 23% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 13% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 12% | 0% |
docs/_generated/active_projects.md.Build a claim inventory first (numbered, file:line, class: number, comparison, citation, cross-ref, method, figure reading).
Pass 1 — Source traceability: For each claim, locate backing artifact (generated variable, src/ + test, output/ file, bib key, cross-ref key). Tag BACKED / WEAK / UNBACKED. The canonical deterministic check for binding manuscript numbers/citations to registered evidence is infrastructure.validation.cli evidence <project> --fail-on-issues — run it as the machine gate for BACKED claims.
Pass 2 — Independent recomputation: Regenerate from clean; re-run validation CLI. Reconciliation table: stated | recomputed | tolerance | PASS/FAIL. UNREPRODUCIBLE = FAIL.
Pass 3 — Methodological adversary: Independent review axis — overstated claims, citation mismatch, figure/caption issues, abstract not earned by results. Tag OVERSTATED / UNSUPPORTED / MISLEADING.
Pass 4 — Reference existence (anti-hallucination): Resolve every cited reference against Crossref / OpenAlex / arXiv with the deterministic verification gate. Tag each ok / mismatch / fabricated / unverifiable / unchecked / anachronism. fabricated, mismatch, and anachronism are blocking. unchecked (offline + uncached) is an honest non-pass — re-run with --live to resolve, never treat it as clean. Distilled from the ARS hallucination taxonomy; runs fully offline against the SQLite cache once seeded.
Improve (not just report): Tighten claims; use generated-variable mechanism; fix citations/cross-refs; keep renderable format ([@key], @fig: OR [[…]] per project — never raw \cite{}/\ref{}). Update projects/<n>/AGENTS.md and README.md. Never hand-edit output/.
fabricated/mismatch/anachronism; unchecked resolved via a --live pass)bashuv sync uv run python scripts/runner/execute_pipeline.py --project <project> --core-only uv run pytest projects/<project>/tests/ --cov=projects/<project>/src --cov-fail-under=90 -q uv run python -m infrastructure.validation.cli prerender projects/<project>/manuscript --repo-root . uv run python -m infrastructure.validation.cli evidence projects/<project> --manuscript-dir projects/<project>/manuscript --fail-on-issues uv run python -m infrastructure.reference.citation validate projects/<project>/manuscript/references.bib uv run python -m infrastructure.reference.verification verify projects/<project>/manuscript/references.bib --live --as-of-year <year> --fail-on-issues uv run python -m infrastructure.validation.cli markdown projects/<project>/manuscript --repo-root . --strict uv run python -m infrastructure.validation.cli links --repo-root . uv run python -m infrastructure.validation.cli pdf output/<project>/pdf/ uv run python -m infrastructure.validation.cli integrity output/<project>/ uv run python -m infrastructure.validation.cli prose-quality projects/<project>/manuscript uv run python -m infrastructure.prose.cli report projects/<project>/manuscript
The reference-existence gate is offline-first: drop --live to verify against the cached resolutions only (CI-safe), and keep --live for the seeding pass that populates the cache. prose-quality is advisory — add --fail-on-flags to gate.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-13 | fail→pass | 18,247 | 1,843 | -90% | 1 | 1 | 0% | 1,608 | 1,426 | -11% | 0 | 0 | — |
case-05 | fail→pass | 12,393 | 7,035 | -43% | 1 | 1 | 0% | 2,079 | 2,110 | +1% | 0 | 0 | — |
case-11 | fail→pass | 9,785 | 5,548 | -43% | 1 | 1 | 0% | 1,728 | 2,132 | +23% | 0 | 0 | — |
case-12 | fail→pass | 7,245 | 3,287 | -55% | 1 | 1 | 0% | 1,373 | 1,552 | +13% | 0 | 0 | — |
case-01 | fail→fail | 4,300 | 6,324 | +47% | 1 | 1 | 0% | 167 | 1,439 | +762% | 0 | 0 | — |
case-02 | fail→fail | 3,872 | 5,586 | +44% | 1 | 1 | 0% | 175 | 1,402 | +701% | 0 | 0 | — |
case-03 | fail→fail | 20,730 | 4,742 | -77% | 1 | 1 | 0% | 4,161 | 1,402 | -66% | 0 | 0 | — |
case-04 | fail→fail | 12,133 | 2,001 | -84% | 1 | 1 | 0% | 1,656 | 1,461 | -12% | 0 | 0 | — |
case-06 | fail→pass | 9,471 | 4,474 | -53% | 1 | 1 | 0% | 1,587 | 1,779 | +12% | 0 | 0 | — |
case-07 | fail→pass | 12,707 | 5,568 | -56% | 1 | 1 | 0% | 2,071 | 2,094 | +1% | 0 | 0 | — |
case-08 | fail→pass | 6,585 | 4,615 | -30% | 1 | 1 | 0% | 1,157 | 1,946 | +68% | 0 | 0 | — |
case-09 | fail→pass | 9,186 | 3,726 | -59% | 1 | 1 | 0% | 1,591 | 1,758 | +10% | 0 | 0 | — |
case-10 | pass→pass | 6,430 | 3,807 | -41% | 1 | 1 | 0% | 1,077 | 1,724 | +60% | 0 | 0 | — |
case-14 | fail→pass | 14,714 | 1,614 | -89% | 1 | 1 | 0% | 2,578 | 1,353 | -48% | 0 | 0 | — |
case-15 | fail→fail | 7,849 | 2,241 | -71% | 1 | 1 | 0% | 1,310 | 1,385 | +6% | 0 | 0 | — |
case-16 | fail→fail | 8,611 | 1,805 | -79% | 1 | 1 | 0% | 1,366 | 1,359 | -1% | 0 | 0 | — |
case-17 | fail→pass | 10,152 | 1,653 | -84% | 1 | 1 | 0% | 1,453 | 1,294 | -11% | 0 | 0 | — |
case-18 | fail→fail | 9,181 | 1,407 | -85% | 1 | 1 | 0% | 1,535 | 1,259 | -18% | 0 | 0 | — |
case-19 | fail→pass | 14,029 | 3,778 | -73% | 1 | 1 | 0% | 2,829 | 1,713 | -39% | 0 | 0 | — |
case-20 | pass→pass | 8,624 | 2,466 | -71% | 1 | 1 | 0% | 1,471 | 1,488 | +1% | 0 | 0 | — |
case-21 | fail→pass | 19,862 | 8,796 | -56% | 1 | 1 | 0% | 2,798 | 2,057 | -26% | 0 | 0 | — |
case-22 | pass→pass | 7,479 | 7,269 | -3% | 1 | 1 | 0% | 1,162 | 2,525 | +117% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 19 counted toward the lift figure. The other 3 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +55 percentage points is the difference between those two pass rates over the 19 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.