Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Write or rewrite a document in evidence-locked mode: no unsourced sentences — every substantive claim carries a footnote citing the exact passage in the user's provided sources, and anything unsupportable is explicitly marked. Use when asked to make a document fully sourced, add citations from my docs, ground a draft in the attached material, or produce something for audiences that will check (legal, board, regulators, enterprise buyers). Produces the document with numbered citations, a source m
.claude/skills/mohitagw15856-evidence-lock/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-10 | ✗→✓ | ▲ Improved | -44% | 0% |
| case-14 | ✗→✓ | ▲ Improved | 139% | 0% |
| case-04 | ✓→✗ | ▼ Worse | 92% | 0% |
| case-07 | ✓→✗ | ▼ Worse | 29% | 0% |
| case-16 | ✓→✗ | ▼ Worse | 132% | 0% |
For most drafts, plausible is enough. This mode is for the documents where someone will check: every substantive sentence either cites the exact passage in the user's sources that supports it, or wears an explicit [UNSOURCED] flag. No third state.
[UNSOURCED] item with what evidence would resolve itAsk for (if not already provided):
[UNSOURCED]). Default: soft.[3] → "Q2 churn analysis, §4: 'logo churn concentrated in accounts under $10k ACV (71% of losses)'". Citing a whole document is not a lock.[UNSOURCED] (or the sentence gets weakened to what the source supports — prefer weakening).[inference from 2,5] — distinguishing sourced, inferred from sourced, and unsourced.The document. Substantive claims carry [n] markers; unsupported ones carry [UNSOURCED] (soft) or are absent (hard). Inferences carry [inference from n,m].]
Source map | # | Source | Supporting passage (verbatim) | |---|---|---| | 1 | doc, section] | "exact quote]" |
Unsupported claims register | Claim | Why it's unsourced | What would resolve it | |---|---|---|
Conflicts noted: source A says X; source B says Y — surfaced at footnote n]
[UNSOURCED] flag, or an [inference from …] label — zero unmarked claims| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 8,520 | 5,502 | -35% | 1 | 1 | 0% | 1,266 | 1,814 | +43% | 0 | 0 | — |
case-02 | fail→fail | 2,532 | 4,017 | +59% | 1 | 1 | 0% | 370 | 1,631 | +341% | 0 | 0 | — |
case-03 | fail→fail | 4,393 | 3,788 | -14% | 1 | 1 | 0% | 616 | 1,613 | +162% | 0 | 0 | — |
case-04 | pass→fail | 5,732 | 5,259 | -8% | 1 | 1 | 0% | 959 | 1,839 | +92% | 0 | 0 | — |
case-05 | pass→pass | 5,807 | 12,033 | +107% | 1 | 1 | 0% | 963 | 2,859 | +197% | 0 | 0 | — |
case-06 | pass→pass | 8,117 | 6,284 | -23% | 1 | 1 | 0% | 1,265 | 2,074 | +64% | 0 | 0 | — |
case-07 | pass→fail | 6,481 | 2,779 | -57% | 1 | 1 | 0% | 1,083 | 1,392 | +29% | 0 | 0 | — |
case-08 | pass→pass | 2,515 | 4,612 | +83% | 1 | 1 | 0% | 375 | 1,701 | +354% | 0 | 0 | — |
case-09 | pass→pass | 5,724 | 10,862 | +90% | 1 | 1 | 0% | 886 | 2,877 | +225% | 0 | 0 | — |
case-10 | fail→pass | 17,631 | 4,219 | -76% | 1 | 1 | 0% | 2,971 | 1,665 | -44% | 0 | 0 | — |
case-11 | fail→fail | 3,292 | 3,379 | +3% | 1 | 1 | 0% | 528 | 1,540 | +192% | 0 | 0 | — |
case-12 | fail→fail | 5,878 | 3,599 | -39% | 1 | 1 | 0% | 817 | 1,511 | +85% | 0 | 0 | — |
case-13 | pass→pass | 4,178 | 5,357 | +28% | 1 | 1 | 0% | 715 | 1,965 | +175% | 0 | 0 | — |
case-14 | fail→pass | 6,596 | 8,840 | +34% | 1 | 1 | 0% | 1,041 | 2,491 | +139% | 0 | 0 | — |
case-15 | pass→pass | 4,797 | 5,391 | +12% | 1 | 1 | 0% | 725 | 1,801 | +148% | 0 | 0 | — |
case-16 | pass→fail | 5,562 | 7,108 | +28% | 1 | 1 | 0% | 905 | 2,104 | +132% | 0 | 0 | — |
case-17 | pass→pass | 4,455 | 2,911 | -35% | 1 | 1 | 0% | 729 | 1,432 | +96% | 0 | 0 | — |
case-18 | fail→fail | 1,416 | 2,922 | +106% | 1 | 1 | 0% | 196 | 1,362 | +595% | 0 | 0 | — |
case-19 | pass→fail | 17,171 | 8,589 | -50% | 1 | 1 | 0% | 2,161 | 2,277 | +5% | 0 | 0 | — |
case-20 | pass→pass | 6,315 | 7,928 | +26% | 1 | 1 | 0% | 1,365 | 2,631 | +93% | 0 | 0 | — |
case-21 | pass→pass | 11,954 | 6,349 | -47% | 1 | 1 | 0% | 1,528 | 1,919 | +26% | 0 | 0 | — |
case-22 | fail→fail | 8,079 | 3,749 | -54% | 1 | 1 | 0% | 1,247 | 1,557 | +25% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of -43 percentage points is the difference between those two pass rates over the 22 comparable cases. 5 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.