Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Audit any document against its own sources — every factual claim extracted and graded as evidenced, partially evidenced, unsupported, or contradicted, with the exact source line that supports or fails it. Use when asked to fact-check a document against its sources, check whether a report's claims are backed up, verify a deck against the data, or ask 'does this doc have receipts?'. Produces a claim ledger, unsupported claims ranked by load-bearingness, a fix-or-drop call per claim, and an honesty
.claude/skills/mohitagw15856-receipts-audit/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-03 | ✗→✓ | ▲ Improved | 70% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 42% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 179% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 41% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 160% | 0% |
A document earns trust one sourced claim at a time. This skill takes a document plus the evidence behind it, extracts every factual claim, and grades each strictly against the provided sources — never against plausibility. The output is a claim ledger, a ranked list of the unsupported claims that matter most, and an honesty score whose method is shown, not asserted.
Ask for these if not provided:
1. Extract claims. Every checkable factual assertion: numbers, comparisons, dates, quotes, causal statements ("X drove Y"), and superlatives ("fastest", "first"). Opinions and hedged forecasts are out of scope — note them separately, don't grade them.
2. Grade each claim against the sources only:
| Grade | Test | |---|---| | Evidenced | A specific source line states it, at the claim's full scope and tense | | Partially evidenced | A source supports a narrower, older, or weaker version than the claim as worded | | Unsupported | No provided source addresses it — even if it is probably true | | Contradicted | A provided source says otherwise; quote both lines side by side |
3. Rank the failures by load-bearingness (1–5): 5 = the document's core conclusion collapses without it; 3 = a supporting pillar; 1 = colour. Rank only Unsupported and Contradicted claims.
4. Fix-or-drop per failing claim. Fix = rewrite so an existing source line covers it (narrow the scope, soften the tense, attribute it). Drop = no source can carry it. Name the evidence that, if obtained, would upgrade it.
5. Score honesty: 100 × (Evidenced + 0.5 × Partially) / total graded claims, then subtract 10 per Contradicted claim, floor 0. State this formula and the counts in the output.
1. Verdict — honesty score, counts per grade, and the single most consequential failing claim.
2. Claim ledger — table: # | claim (verbatim) | grade | source line (quoted, with location) or "none provided".
3. Load-bearing failures — failing claims ranked 5→1, each with one sentence on what breaks if it's wrong.
4. Fix-or-drop register — per failing claim: Fix (rewritten wording + the source line it now rests on) or Drop (why no fix exists), plus evidence-to-collect.
5. Honesty score method — the formula, the counts, the arithmetic.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-03 | fail→pass | 34,728 | 127,347 | +267% | 1 | 1 | 0% | 2,318 | 3,949 | +70% | 0 | 0 | — |
case-04 | pass→pass | 18,267 | 15,713 | -14% | 1 | 1 | 0% | 1,799 | 4,214 | +134% | 0 | 0 | — |
case-01 | fail→fail | 12,107 | 9,980 | -18% | 1 | 1 | 0% | 1,060 | 1,836 | +73% | 0 | 0 | — |
case-02 | fail→pass | 18,924 | 18,254 | -4% | 1 | 1 | 0% | 2,381 | 3,385 | +42% | 0 | 0 | — |
case-05 | pass→pass | 9,152 | 10,764 | +18% | 1 | 1 | 0% | 1,382 | 2,997 | +117% | 0 | 0 | — |
case-06 | pass→pass | 9,755 | 29,796 | +205% | 1 | 1 | 0% | 1,392 | 3,348 | +141% | 0 | 0 | — |
case-07 | fail→pass | 9,043 | 18,808 | +108% | 1 | 1 | 0% | 1,313 | 3,669 | +179% | 0 | 0 | — |
case-08 | fail→pass | 14,917 | 11,381 | -24% | 1 | 1 | 0% | 2,177 | 3,078 | +41% | 0 | 0 | — |
case-09 | pass→pass | 7,654 | 13,449 | +76% | 1 | 1 | 0% | 1,257 | 2,972 | +136% | 0 | 0 | — |
case-10 | fail→pass | 13,848 | 14,504 | +5% | 1 | 1 | 0% | 1,332 | 3,457 | +160% | 0 | 0 | — |
case-11 | fail→fail | 6,536 | 9,580 | +47% | 1 | 1 | 0% | 963 | 2,423 | +152% | 0 | 0 | — |
case-12 | pass→fail | 27,378 | 6,258 | -77% | 1 | 1 | 0% | 3,427 | 2,052 | -40% | 0 | 0 | — |
case-13 | pass→pass | 7,095 | 20,482 | +189% | 1 | 1 | 0% | 1,264 | 3,340 | +164% | 0 | 0 | — |
case-14 | pass→pass | 11,643 | 16,093 | +38% | 1 | 1 | 0% | 1,634 | 3,357 | +105% | 0 | 0 | — |
case-15 | pass→fail | 11,338 | 20,584 | +82% | 1 | 1 | 0% | 1,817 | 3,973 | +119% | 0 | 0 | — |
case-16 | pass→pass | 11,168 | 6,159 | -45% | 1 | 1 | 0% | 1,602 | 2,036 | +27% | 0 | 0 | — |
case-17 | pass→pass | 6,963 | 14,008 | +101% | 1 | 1 | 0% | 1,327 | 3,110 | +134% | 0 | 0 | — |
case-18 | pass→pass | 16,743 | 31,483 | +88% | 1 | 1 | 0% | 2,150 | 4,163 | +94% | 0 | 0 | — |
case-19 | fail→pass | 9,677 | 5,412 | -44% | 1 | 1 | 0% | 1,469 | 1,909 | +30% | 0 | 0 | — |
case-20 | pass→pass | 10,913 | 14,375 | +32% | 1 | 1 | 0% | 1,533 | 3,104 | +102% | 0 | 0 | — |
case-21 | pass→pass | 9,480 | 6,757 | -29% | 1 | 1 | 0% | 1,347 | 2,212 | +64% | 0 | 0 | — |
case-22 | pass→pass | 26,159 | 6,497 | -75% | 1 | 1 | 0% | 2,207 | 1,986 | -10% | 0 | 0 | — |
case-23 | fail→pass | 14,438 | 23,190 | +61% | 1 | 1 | 0% | 2,096 | 4,374 | +109% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of +22 percentage points is the difference between those two pass rates over the 23 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.