Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Qualify a check before it is allowed to clear anything — prove it can detect the failure it is meant to catch. Seeds known defects into a copy of a real artifact plus a clean control, runs the checker, and reports recall and false-positive rate into a qualification ledger. Use when the user says "does this check work", "qualify the gate", "test my reviewer", "seed defects", "vaccinate", "qualify the checks", "can I trust this review", "does it pass for the right reason", or before relying on any
.claude/skills/pedrohcgs-vaccinate/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-07 | ✗→✓ | ▲ Improved | 124% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 61% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 51% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 17% | 0% |
| case-16 | ✗→✓ | ▲ Improved | 116% | 0% |
Twenty bugs were once planted in a working codebase and the review agents were asked to check it again. They reported everything was fine. Recall: 0/20.
A vaccine is a small, controlled dose of error that strengthens the whole system. This skill administers one.
The rule it enforces: an unqualified check is not weak evidence — it is none.
decision that matters (a submission, a release, a deposit).
State the defect class the check is supposed to catch. "Catches problems" is not a class. "Detects a coefficient in the text that no longer matches its table" is.
Work on a copy, never the live artifact. Produce:
references/defect-library.md.The control is not optional. Without it you measure recall and call it accuracy.
Verify each seed actually violates something. A seed that the artifact already permits creates no defect, and the checker correctly reporting "pass" will look like a broken gate. This is the most common way a qualification run produces a false alarm about itself.
Run the check or agent against each variant in a fresh context, one variant per run. It must not know which variant it has, how many defects exist, or that a qualification is underway. For an AI reviewer, spawn via the Agent tool with context: fork.
| Metric | Definition | |---|---| | Recall | seeded defects correctly identified / seeded defects planted | | False-positive rate | findings on the clean control that are factually false / total findings on the control | | Localization | did it name the right location, or just report unease? | | Baseline delta | recall of a simpler alternative (a grep, a diff, a one-line assertion) |
A finding on the clean control counts as a false positive only when it is factually wrong — not merely unwelcome. A reviewer prompted to find gaps will report some in sound work; that is expected behaviour, not a failure.
The baseline is load-bearing. A five-agent panel that scores no better than grep -n has not earned its cost.
Append to quality_reports/qualification/LEDGER.md:
| date | target | artifact | defect classes | N | recall | FPR | baseline | verdict |Verdicts: PASS (detects its named class at an agreed threshold) · FAIL (misses it) · BLOCKED (could not be run — say why; do not record as PASS).
weaken the seed until it passes.
nothing.
/vaccinate check-model-versions.shThe newest model is Opus 4.8 and it is the default. to README.md.Control: unmodified README.md.
bash scripts/check-model-versions.sh; echo $?Recall 1/1, FPR 0/0. Baseline: grep -c "Opus 4.8" README.md also detects — so the gate's value is its allow-marker logic, not raw detection.
PASS.broken gate.
replicates per class where cost allows.
| File | Read when | |---|---| | references/defect-library.md | choosing what to seed — defect classes by artifact type | | evals/README.md | the complementary question: does the skill produce better output than not having it? |
A second model, more agents, or a longer debate is not presumed to verify better. Before an elaborate procedure earns extra weight, show it outperforms a simpler check on the same prespecified seeded failures and valid cases, reporting both detection and false alarms. Complexity that has not beaten a baseline is cost, not assurance.
When a model grades, triages, or reviews at scale:
Any material change to the model, prompt, rubric, or target population requires fresh human labels and recalibration.
A check qualified against an old interface, schema, or scale may silently stop testing anything. Re-run the seeded-defect proof after material changes to the object under test or to the check itself.
regenerates, an estimator recovers an analytic special case, a seeded fault triggers a failure. These can be automated and rerun forever.
identifying assumption is plausible, whether a result deserves causal language — cannot be automated, and no volume of qualified checks substitutes for one.
A missing, substituted, or degraded check is missing evidence, not a pass. Verify the run happened (log, exit status, artifact timestamp — not an assumption); that it ran on the current object, not a cached one; that nothing was skipped, filtered, or swallowed into a default; and that the tolerance was fixed before the comparison. A tolerance loosened after a failed comparison converts evidence into decoration. If it must be loosened, record it as an approved divergence with a reason.
verification-ladder.md — rung 0; why this comes before everything/qualify-checks (2026-08-21): same goal — one skill, not twoexternal-oracle-process.md — qualifying an external referee| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 8,866 | 10,927 | +23% | 1 | 1 | 0% | 426 | 2,172 | +410% | 0 | 0 | — |
case-02 | fail→fail | 7,498 | 11,019 | +47% | 1 | 1 | 0% | 262 | 2,464 | +840% | 0 | 0 | — |
case-03 | fail→fail | 6,371 | 9,509 | +49% | 1 | 1 | 0% | 170 | 2,052 | +1107% | 0 | 0 | — |
case-04 | pass→pass | 18,234 | 12,099 | -34% | 1 | 1 | 0% | 3,120 | 3,866 | +24% | 0 | 0 | — |
case-05 | pass→fail | 15,206 | 16,069 | +6% | 1 | 1 | 0% | 2,674 | 2,104 | -21% | 0 | 0 | — |
case-06 | pass→fail | 73,393 | 24,714 | -66% | 1 | 1 | 0% | 2,438 | 4,854 | +99% | 0 | 0 | — |
case-07 | fail→pass | 40,285 | 15,179 | -62% | 1 | 1 | 0% | 1,711 | 3,838 | +124% | 0 | 0 | — |
case-08 | fail→fail | 25,831 | 3,754 | -85% | 1 | 1 | 0% | 1,095 | 2,293 | +109% | 0 | 0 | — |
case-09 | pass→pass | 14,011 | 9,338 | -33% | 1 | 1 | 0% | 1,926 | 2,953 | +53% | 0 | 0 | — |
case-10 | fail→pass | 63,122 | 21,809 | -65% | 1 | 1 | 0% | 2,215 | 3,565 | +61% | 0 | 0 | — |
case-11 | fail→pass | 11,643 | 4,840 | -58% | 1 | 1 | 0% | 1,639 | 2,472 | +51% | 0 | 0 | — |
case-12 | fail→pass | 48,810 | 12,130 | -75% | 1 | 1 | 0% | 2,214 | 2,586 | +17% | 0 | 0 | — |
case-13 | pass→pass | 16,516 | 11,229 | -32% | 1 | 1 | 0% | 2,334 | 3,256 | +40% | 0 | 0 | — |
case-14 | pass→pass | 20,683 | 13,308 | -36% | 1 | 1 | 0% | 1,741 | 2,739 | +57% | 0 | 0 | — |
case-15 | pass→pass | 73,470 | 13,668 | -81% | 1 | 1 | 0% | 3,001 | 3,503 | +17% | 0 | 0 | — |
case-16 | fail→pass | 8,479 | 25,145 | +197% | 1 | 1 | 0% | 1,286 | 2,777 | +116% | 0 | 0 | — |
case-17 | fail→fail | 13,783 | 12,710 | -8% | 1 | 1 | 0% | 1,958 | 3,575 | +83% | 0 | 0 | — |
case-18 | pass→pass | 45,136 | 10,619 | -76% | 1 | 1 | 0% | 1,971 | 3,290 | +67% | 0 | 0 | — |
case-19 | fail→pass | 53,973 | 12,107 | -78% | 1 | 1 | 0% | 1,123 | 2,972 | +165% | 0 | 0 | — |
case-20 | fail→fail | 17,983 | 33,398 | +86% | 1 | 1 | 0% | 2,430 | 3,729 | +53% | 0 | 0 | — |
case-21 | pass→pass | 10,214 | 7,056 | -31% | 1 | 1 | 0% | 1,523 | 2,707 | +78% | 0 | 0 | — |
case-22 | fail→pass | 38,657 | 7,910 | -80% | 1 | 1 | 0% | 1,740 | 2,573 | +48% | 0 | 0 | — |
case-23 | pass→pass | 9,668 | 12,912 | +34% | 1 | 1 | 0% | 1,452 | 2,575 | +77% | 0 | 0 | — |
case-24 | fail→pass | 56,281 | 10,755 | -81% | 1 | 1 | 0% | 1,875 | 3,293 | +76% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 24 cases were attempted, and 18 counted toward the lift figure. The other 6 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +25 percentage points is the difference between those two pass rates over the 18 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.