Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Force verification before answering questions involving versions, dates, status, or "current" state. Prevents hallucinations at the REASONING level by checking assertions.yaml and WebSearch before forming beliefs. Triggers on software versions, release status, dates, and package versions.
.claude/skills/aiskillstore-epistemic-checkpoint/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-12 | ✗→✓ | ▲ Improved | 65% | 0% |
| case-02 | ✗→✓ | ▲ Improved | -6% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 38% | 0% |
| case-04 | ✗→✓ | ▲ Improved | -29% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 30% | 0% |
Force verification before answering questions involving versions, dates, status, or "current" state.
Prevents the ROOT CAUSE of hallucinations - not just blocking wrong output, but preventing wrong REASONING. Claude's training data is stale; this skill forces verification before forming beliefs.
Activate this skill when the question involves ANY of:
Say to yourself: "My training data may be stale for: topic]"
textRead ${CLAUDE_PLUGIN_ROOT}/blackboard/assertions.yaml Search for relevant entries
If found - use that value with high confidence.
textWebSearch("[software] [version] release date site:official-docs") WebSearch("[software] LTS release 2025")
Prefer official sources:
Say: "Based on source], software] version] is status] as of date]."
Only now answer the actual question with the verified baseline.
| If you think... | Actually do... | |-----------------|----------------| | "I'm pretty sure .NET 10 is..." | WebSearch to verify | | "This is probably still preview" | Check assertions.yaml | | "I remember this from training" | Training is stale, verify | | "This is a simple factual question" | Simple facts are often WRONG | | "The user said it's preview" | User might be wrong too, verify |
WRONG:
text"If targeting .NET 10 preview, use C# 14 extension types..."
CORRECT:
text"Let me verify .NET 10 status... [WebSearch] ... .NET 10 is LTS (released November 2025), not preview. Standard extension methods work fine."
WRONG:
text"React 19 is still in beta, so..."
CORRECT:
text"Checking React 19 status... [WebSearch] ... React 19 was released December 2024 as stable. Proceeding with React 19 stable patterns."
| Topic | Danger Pattern | Likely Truth | |-------|----------------|--------------| | .NET | "preview", "not released" | .NET 10 LTS since Nov 2025 | | React | "beta", "experimental" | React 19 stable since Dec 2024 | | Node | "current" without version | Always specify exact version | | Python | "3.x is latest" | Verify - 3.13+ exists |
When this skill activates, structure your verification as:
textEpistemic Checkpoint Claim to verify: [what you were about to assume] Source checked: [assertions.yaml / WebSearch / official docs] Verified fact: [the actual truth] Confidence: [high/medium/low] Proceeding with verified baseline...
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-11 | pass→pass | 17,717 | 27,290 | +54% | 1 | 1 | 0% | 2,164 | 2,950 | +36% | 0 | 0 | — |
case-12 | fail→pass | 10,818 | 20,515 | +90% | 1 | 1 | 0% | 1,862 | 3,073 | +65% | 0 | 0 | — |
case-13 | pass→pass | 11,169 | 21,808 | +95% | 1 | 1 | 0% | 2,042 | 2,062 | +1% | 0 | 0 | — |
case-22 | pass→pass | 4,092 | 3,342 | -18% | 1 | 1 | 0% | 784 | 1,339 | +71% | 0 | 0 | — |
case-01 | fail→fail | 15,482 | 16,459 | +6% | 1 | 1 | 0% | 1,583 | 1,135 | -28% | 0 | 0 | — |
case-02 | fail→pass | 19,732 | 24,341 | +23% | 1 | 1 | 0% | 3,725 | 3,498 | -6% | 0 | 0 | — |
case-03 | fail→pass | 12,693 | 12,577 | -1% | 1 | 1 | 0% | 2,432 | 3,344 | +38% | 0 | 0 | — |
case-04 | fail→pass | 21,647 | 17,008 | -21% | 1 | 1 | 0% | 3,342 | 2,380 | -29% | 0 | 0 | — |
case-05 | fail→pass | 20,554 | 14,692 | -29% | 1 | 1 | 0% | 3,013 | 3,904 | +30% | 0 | 0 | — |
case-06 | fail→pass | 12,170 | 11,346 | -7% | 1 | 1 | 0% | 1,796 | 2,516 | +40% | 0 | 0 | — |
case-07 | fail→fail | 18,419 | 13,374 | -27% | 1 | 1 | 0% | 2,209 | 2,273 | +3% | 0 | 0 | — |
case-08 | pass→pass | 14,836 | 13,500 | -9% | 1 | 1 | 0% | 2,062 | 2,441 | +18% | 0 | 0 | — |
case-09 | pass→pass | 13,337 | 13,708 | +3% | 1 | 1 | 0% | 1,352 | 2,303 | +70% | 0 | 0 | — |
case-10 | pass→pass | 5,628 | 8,729 | +55% | 1 | 1 | 0% | 859 | 1,450 | +69% | 0 | 0 | — |
case-14 | pass→pass | 9,372 | 5,677 | -39% | 1 | 1 | 0% | 1,518 | 1,702 | +12% | 0 | 0 | — |
case-15 | pass→pass | 6,692 | 6,286 | -6% | 1 | 1 | 0% | 1,107 | 1,825 | +65% | 0 | 0 | — |
case-16 | pass→pass | 8,624 | 15,681 | +82% | 1 | 1 | 0% | 1,379 | 2,660 | +93% | 0 | 0 | — |
case-17 | fail→fail | 7,044 | 10,553 | +50% | 1 | 1 | 0% | 1,249 | 1,042 | -17% | 0 | 0 | — |
case-18 | fail→pass | 15,427 | 6,866 | -55% | 1 | 1 | 0% | 1,576 | 1,111 | -30% | 0 | 0 | — |
case-19 | fail→pass | 9,741 | 5,459 | -44% | 1 | 1 | 0% | 884 | 1,739 | +97% | 0 | 0 | — |
case-20 | pass→pass | 11,140 | 8,002 | -28% | 1 | 1 | 0% | 1,106 | 2,418 | +119% | 0 | 0 | — |
case-21 | pass→pass | 5,264 | 4,032 | -23% | 1 | 1 | 0% | 848 | 1,472 | +74% | 0 | 0 | — |
case-23 | pass→pass | 11,175 | 12,863 | +15% | 1 | 1 | 0% | 1,947 | 2,171 | +12% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 21 counted toward the lift figure. The other 2 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +35 percentage points is the difference between those two pass rates over the 21 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.