Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Prove that the file you are about to send, publish, or hand off IS the thing you mean — before it leaves your hands. Rebuild from source, check integrity, diff the derived artifact against its source, and require the recipient to echo what they received. Use before sending a paper/PDF for review, shipping a release or replication package, handing files to a coauthor or collaborator, uploading anything to an external reviewer or model, or publishing. Prevents an entire review/QA cycle from being
.claude/skills/pedrohcgs-verify-artifact/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 112% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 36% | 0% |
| case-14 | ✗→✓ | ▲ Improved | 32% | 0% |
| case-15 | ✗→✓ | ▲ Improved | 79% | 0% |
| case-17 | ✗→✓ | ▲ Improved | 39% | 0% |
A review is only as good as the thing reviewed. If the artifact is corrupt, truncated, stale, or silently renumbered, a competent reviewer will return confident findings about defects that do not exist in your work — and you will spend a cycle chasing them. This is cheap to prevent and expensive to miss.
Rule: never send a derived artifact you have not diffed against its source.
Never send yesterday's build. Rebuild, then copy to the send location in the same command, so nothing can change underneath you.
Cloud-synced folders (Dropbox/iCloud/OneDrive) dehydrate files: a PDF can become 0 bytes or a partial copy between building and reading it. Symptoms: Syntax Error: Couldn't find trailer dictionary, a 70 KB file that should be 700 KB. Build-tool "up to date" messages are not evidence the file is intact — the tool checks timestamps, not content. Always stage to a local (non-synced) directory and verify there.
pdfinfo, pdftotext, a JSON/CSV parser, unzip -t). A file that exists is not a file that works.??, [cite], TODO, XXX, \ref{ leftovers, "Chapter ??"; for code/data, NaNs, empty cells, placeholder values.This is the step people skip and the one that pays. If you produced the artifact by transforming, excerpting, compressing, or subsetting, then enumerate what could have been lost and check it:
If you cut anything, leave a visible in-artifact note saying so, so a reviewer does not read an omission as a gap.
Ask the reviewer (human or model) to state, at the top of their response, the exact filenames, page/record counts, and version they are reviewing. This catches stale caches, wrong attachments, and silent fallbacks to an older upload — failures that are otherwise invisible until the findings make no sense.
Use unique filenames per round (report_r6.pdf, not report.pdf). Repeated identical names invite the recipient's system to serve a cached earlier copy.
Before acting on a review, ask: could this finding be an artifact of what I sent? Signals: complaints about missing/undefined references, "sections appear truncated", "the proof ends mid-argument", numbering that does not match your copy, or objections to text you know is present. Re-verify the artifact before you re-verify the work. Applying "fixes" for artifact-induced findings actively damages correct material.
verification-ladder.md — rung 2 (existence → substantiveness → wiring → coherence)external-oracle-process.md §7 — cloud-synced files upload corrupt| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 17,663 | 34,195 | +94% | 1 | 1 | 0% | 1,860 | 3,935 | +112% | 0 | 0 | — |
case-02 | fail→fail | 15,352 | 20,101 | +31% | 1 | 1 | 0% | 2,736 | 4,208 | +54% | 0 | 0 | — |
case-03 | fail→fail | 13,268 | 29,625 | +123% | 1 | 1 | 0% | 377 | 4,115 | +992% | 0 | 0 | — |
case-04 | fail→fail | 11,837 | 8,292 | -30% | 1 | 1 | 0% | 1,324 | 2,067 | +56% | 0 | 0 | — |
case-05 | pass→pass | 16,965 | 18,640 | +10% | 1 | 1 | 0% | 2,977 | 3,720 | +25% | 0 | 0 | — |
case-06 | fail→fail | 7,828 | 3,878 | -50% | 1 | 1 | 0% | 1,069 | 1,552 | +45% | 0 | 0 | — |
case-07 | fail→fail | 15,321 | 20,963 | +37% | 1 | 1 | 0% | 2,356 | 2,700 | +15% | 0 | 0 | — |
case-08 | pass→pass | 13,647 | 8,962 | -34% | 1 | 1 | 0% | 1,599 | 2,355 | +47% | 0 | 0 | — |
case-09 | pass→pass | 13,800 | 10,982 | -20% | 1 | 1 | 0% | 1,798 | 2,691 | +50% | 0 | 0 | — |
case-10 | pass→pass | 26,407 | 15,086 | -43% | 1 | 1 | 0% | 2,576 | 3,342 | +30% | 0 | 0 | — |
case-11 | fail→pass | 10,171 | 6,788 | -33% | 1 | 1 | 0% | 1,429 | 1,940 | +36% | 0 | 0 | — |
case-12 | pass→pass | 10,332 | 9,300 | -10% | 1 | 1 | 0% | 1,431 | 2,207 | +54% | 0 | 0 | — |
case-13 | pass→pass | 21,910 | 19,205 | -12% | 1 | 1 | 0% | 2,070 | 2,146 | +4% | 0 | 0 | — |
case-14 | fail→pass | 43,713 | 13,631 | -69% | 1 | 1 | 0% | 2,443 | 3,226 | +32% | 0 | 0 | — |
case-15 | fail→pass | 12,194 | 13,738 | +13% | 1 | 1 | 0% | 1,787 | 3,190 | +79% | 0 | 0 | — |
case-16 | pass→pass | 13,176 | 16,691 | +27% | 1 | 1 | 0% | 1,854 | 2,554 | +38% | 0 | 0 | — |
case-17 | fail→pass | 10,291 | 18,495 | +80% | 1 | 1 | 0% | 1,578 | 2,188 | +39% | 0 | 0 | — |
case-18 | fail→pass | 11,727 | 8,619 | -27% | 1 | 1 | 0% | 1,870 | 2,292 | +23% | 0 | 0 | — |
case-19 | pass→pass | 18,126 | 36,124 | +99% | 1 | 1 | 0% | 2,999 | 2,833 | -6% | 0 | 0 | — |
case-20 | pass→pass | 14,165 | 15,614 | +10% | 1 | 1 | 0% | 2,061 | 2,181 | +6% | 0 | 0 | — |
case-21 | pass→pass | 18,521 | 17,641 | -5% | 1 | 1 | 0% | 2,782 | 3,519 | +26% | 0 | 0 | — |
case-22 | fail→pass | 12,210 | 13,066 | +7% | 1 | 1 | 0% | 1,780 | 2,929 | +65% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 21 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +32 percentage points is the difference between those two pass rates over the 21 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.